Automation exposure indices measure the technical possibility that a technology could perform work — not whether it will, or whether anyone loses a job. Frey and Osborne put 47 percent of US employment at high risk by scoring whole occupations. Arntz and colleagues got 9 percent by scoring tasks instead. Eloundou and colleagues put 19 percent of workers at more than half their tasks exposed to language models. None of them is a forecast.
Three studies, three answers, one question
Three respected studies answered the same question and got 47 percent, 19 percent and 9 percent. The gap between them is the finding rather than a problem with any one of them. The disagreement tells you more than any single figure.
Frey and Osborne scored 702 whole occupations for computerizability and put 47 percent of United States employment at high risk. Arntz, Gregory and Zierahn rebuilt the same question around tasks and got 9 percent across OECD countries. Same question, different unit, an entirely different answer.
Eloundou and colleagues measured exposure to large language models and found 19 percent of workers with more than half their tasks affected. Those are not three estimates of different things; they are three attempts at the same question differing by a factor of five. A factor of five is not a margin of error.
Why the unit of analysis decides everything
Scoring an occupation treats everybody holding that title as doing identical work. That assumption is doing more work than any of the technical judgments in the study. One assumption quietly determines most of the result.
Scoring tasks recognizes that two people with the same job title spend their weeks differently, and that most jobs are bundles in which some parts are automatable and others are not. The difference between those two framings is the difference between the estimates. Methodology rather than new evidence explains the whole gap.
When the analysis drops to tasks and allows for that variation, high-risk estimates fall from the high forties into single digits. That is not a refinement at the margin; it is most of the answer, and it is why the field settled on task-level decomposition as necessary rather than optional. The correction has been accepted for years now.
What each one actually measured
Frey and Osborne applied expert judgment to whether an entire occupation could be computerized. The results came out bimodal, with occupations clustering at the extremes, which is itself a symptom of a method too coarse for the question being asked. Bimodal results usually mean the instrument is blunt.
Arntz, Gregory and Zierahn rebuilt the same underlying question around the task composition of jobs, across a set of OECD countries with comparable survey data. Same question, finer unit, different answer. Comparable survey data made the rebuild possible at all.
Eloundou and colleagues defined exposure precisely and narrowly: whether a language model could cut the time a task takes by at least half while preserving quality. That is a claim about speed rather than about replacement, and the distinction matters enormously. Speed and headcount are related through decisions nobody has made yet.
They do not even agree on the ranking
Comparisons of these indices find that some correlate with each other and some barely do at all. That is a stronger criticism than any disagreement about a headline percentage. Disagreeing on rank is worse than disagreeing on level.
Measures built on machine-learning suitability and measures built on expert occupation scoring have shown little correlation, despite claiming to describe the same underlying phenomenon. They are not producing different estimates of one quantity so much as measuring different things. Two instruments pointing at different phenomena will never reconcile.
That is worth sitting with rather than skipping past. If two indices rank the same occupations differently, at most one of them is describing reality, and there is no external test that tells you which. Nothing in the world adjudicates between the two rankings.
The criticism worth carrying with you
The occupation-level scores track education level closely, which is a suspicious property for a measure of technical automatability. Nothing about a machine’s capability should depend on how long the incumbent studied. Machines are indifferent to the credentials of the person replaced.
That raises an uncomfortable question about what is actually being measured. Are these indices capturing automatability, or are they capturing credential level and calling it automatability? That would be a measurement of something else entirely.
If it is partly the second, then using such a score to advise an individual amounts to telling them their qualifications predict their risk. That is both less useful and less true than it sounds. Advice built on it would be advice about education instead.
Why the numbers get quoted anyway
Forty-seven percent is a headline and nine percent is not. That asymmetry explains most of what happened to these figures in public circulation. Headline value and accuracy are entirely unrelated properties.
The high estimate came first, arrived with a striking number, and entered general circulation before the methodological response was published. By the time the correction existed, the original had been repeated everywhere. First-mover advantage applies to statistics just as it does elsewhere.
You will still see it quoted without the correction, frequently in serious places. That is not dishonesty so much as the ordinary lag between a finding and its critique, and the critique never travels as far as the finding did. Corrections are structurally quieter than the claims they correct.
How to read any exposure figure
Three questions make any published figure usable, and all three have short answers. Ask them before accepting the number into an argument. Each has a short answer available in the abstract.
Is it occupation-level or task-level, which is the single largest determinant of the result. Exposure to what specifically, meaning robotics, general software or language models. And over what horizon, since a claim about a decade and a claim about three years are different claims.
A number without those three attached cannot support a decision about your own work. With them attached it becomes usable, and usually considerably more modest than the headline implied. Qualified figures are less quotable and considerably more useful.
Exposure is not displacement
Every one of these studies measures whether a task could technically be performed by a machine. That is the ceiling rather than the expectation. Ceilings and expectations are different kinds of claim.
None of them measures whether anyone will build it, buy it, or be allowed to deploy it. None accounts for what happens to demand when the work becomes cheaper, which historically has cut both ways. Cheaper services have expanded employment as often as reduced it.
That gap is why exposure figures have consistently run ahead of observed job losses. Automation of a task inside a job is far more common than automation of the job itself. Tasks leave and job titles mostly stay where they were.
What the studies themselves say about this
The researchers are generally clearer about their limits than the coverage of their work is. The papers describe exposure and the headlines describe unemployment. Those are not the same claim by any reading.
Reading the abstract of any study you see quoted takes about ten minutes and usually reveals a narrower claim than the one being attributed to it. The definition of exposure is stated in the first few paragraphs of each. Nothing about the papers is hidden or hard to find.
That is the cheapest quality check available on this entire subject. It also tends to make the research more interesting rather than less, because the actual findings are more specific than the summaries. Specificity is what makes them usable for a decision.
Common questions
What does an exposure score measure?
The technical possibility that a technology could perform work. Not whether employers will adopt it, and not job losses.
Why is 47 percent so different from 9 percent?
Unit of analysis. Occupation-level scoring treats every holder of a title identically; task-level scoring accounts for the fact that they do different work.
What did the 19 percent figure measure?
The share of workers with more than half their tasks exposed to language models, where exposure means cutting task time by at least half without losing quality.
Do the indices agree with each other?
Not reliably. Some correlate and some barely do, which means at most one of any disagreeing pair describes reality.
How should I read an exposure number?
Ask whether it is occupation or task level, exposure to what exactly, and over what horizon. Without those three it cannot support a decision.
How much employment is at risk from automation?
Published estimates range from 47 percent to 9 percent depending on whether whole occupations or individual tasks are scored. There is no official figure.
Why is the range so wide?
Unit of analysis. Occupation-level scoring treats everyone with a job title identically; task-level scoring accounts for the fact that they do different work.
Do the indices agree with each other?
Not reliably. Some correlate and some barely do, which means at most one of any disagreeing pair describes reality — and no test tells you which.