287 STEM occupations · 5,717 tasks · 963 subtasks

Two measures of automation risk. The same jobs. No relationship at all.

Frey and Osborne put mathematicians in their lowest-risk band in 2013 — a 4.7 per cent chance of computerisation. Score the same occupations for exposure to large language models and mathematicians ranks third of 268. Across the 150 jobs carrying both measures the correlation is 0.006. Not weak agreement. None.

Scored by claude-opus-5
Validated against human expert ratings at r = 0.85

This is what the data says. Every number below comes from a pipeline that scrapes O*NET, scores the 963 distinct work activities underneath 5,717 STEM tasks, propagates those scores back out to every task and occupation, joins BLS employment, and checks itself against published human ratings. The dashboard alongside it lets you interrogate the same tables directly.

01 / The inversion

The jobs most at risk in 2013 are not the jobs most at risk now.

Frey and Osborne scored every US occupation for its probability of computerisation. Score the same occupations for exposure to large language models and the ordering does not weaken — it comes apart entirely.

Fig. 01 — Two measures, same occupations 01 / 03
Mathematicians, 2013 4.7% chance
Left axis square-root scaled; right axis linear. 150 occupations.
01 / 2013

Automation meant routine and manual

Frey and Osborne built their measure on the bottlenecks of that era: perception and manipulation, creative intelligence, social intelligence. Repetitive physical work scored high. Analytical work scored low — mathematicians landed in their lowest-risk band.

Mathematicians — 4.7% probability of computerisation

02 / 2026

Language models run the other way

Rate the same occupations for what a current model could do and the analytical professions move to the top. Mathematicians ranks third of 268. Nothing about the work changed; the capability arriving to meet it did.

Mathematicians — 3rd most exposed of 268

03 / The disagreement

And the old risk ordering does not survive

The occupations Frey and Osborne rated most at risk — surveying technicians at 96 per cent, pharmacy technicians at 92 — sit in the middle of the LLM ranking. Across all 150 the correlation is 0.006. A measure written before 2020 tells you close to nothing about exposure now.

Pearson r = 0.006 · Spearman 0.119

02 / The vocabulary

Five thousand tasks stand on a thousand shared activities.

O*NET decomposes each occupation into tasks, and each task into standardised detailed work activities. The activities repeat across jobs, which is what makes the corpus tractable to score at all.

Fig. 02 — Occupations, tasks, subtasks 01 / 03
Occupations 287
Counts are exact. Marks are illustrative.
01 / The corpus

287 STEM occupations, 5,717 tasks

Every occupation O*NET flags as STEM, scraped with its full task list — each task carrying an importance rating and a Core, Supplemental or New label from the incumbent survey.

5,717 task statements

02 / The shared layer

Only 963 distinct activities underneath

Those tasks map onto a much smaller vocabulary of detailed work activities. “Analyse data to inform operational decisions” is one activity appearing in dozens of different jobs.

963 distinct subtasks · 7,215 links

03 / Why it matters

Score the vocabulary, not the tasks

Rating 963 activities instead of 5,717 tasks is six times less work, and it removes an artefact: the same activity cannot be scored one way in nursing and another in engineering. Consistency comes from the structure rather than from discipline.

Every task inherits the score of its activities

03 / Inside a job

A job is a bundle of tasks, and the bundles differ.

Asking whether an occupation is exposed hides the thing that matters. Almost every job has some tasks a model could take and some it could not. What separates one job from another is the mix.

Fig. 03 — Every task, three jobs 01 / 04
A whole job 100% of tasks
One bar per task. Length is susceptibility.
01 / All of it

Some jobs are exposed all the way down

Business intelligence analysts have seventeen tasks in O*NET. Every one of them scores above the high-exposure threshold. There is no part of the documented job that sits outside it — which is rare.

100% of 17 tasks · spread ±2.4

02 / Part of it

Most jobs split, and the split is the story

Naturopathic physicians have the widest internal spread in the corpus. Their record-keeping, literature review and treatment-planning tasks sit near the top of the exposure range; examining a patient and administering care sit near the bottom. The job does not vanish. It loses one half and keeps the other.

Widest internal spread · 30% of 20 tasks

03 / Almost none of it

And some barely move

Prosthodontists, chemists and medical laboratory technicians have no task above the threshold at all. Not because the work is simple — it is among the most skilled in the corpus — but because the documented tasks are chairside, bench and instrument work.

0% of tasks above threshold

04 / The field

Most jobs lose some tasks, not all

Across 268 scored occupations, 121 have fewer than a fifth of their tasks highly exposed. Eleven have more than four fifths. The common case is partial — a job reshaped around what is left, not one that disappears.

121 under 20% · 11 over 80%

Interlude / Work it through

Pick a job. Decide what counts as exposed.

Every task of every scored occupation, with the score it received. Move the threshold to set how capable you think a model has to be before a task is genuinely at risk — the share of the job it covers moves with you. The ranking of occupations is not fixed; it depends on where you draw that line.

Share of this job

Against all 268 occupations

04 / The shared spine

A few activities run through almost every job.

The subtask vocabulary is not evenly used. Most activities belong to one occupation. A small number appear everywhere — and those are disproportionately the exposed ones.

Fig. 04 — Reach against susceptibility 01 / 04
Used by one job only 261 of 963
Reach axis square-root scaled. 963 distinct activities.
01 / The tail

Most activities belong to a single job

Two hundred and sixty-one of the 963 activities — more than a quarter — appear in exactly one occupation. These are the specialised core of a profession, and automating one of them changes that profession and nothing else.

261 of 963 used by one job

02 / The spine

A handful appear almost everywhere

“Record patient medical histories” appears in 56 different occupations. “Prepare scientific or technical reports” in 40. “Train medical providers” in 49. These are the connective tissue of STEM work rather than anyone’s speciality.

Widest reach — 56 occupations

03 / The overlap

And the shared ones are the exposed ones

The activities with the widest reach are documentation, reporting, literature review and grant writing — exactly the work that scores highest. “Research topics in area of expertise” reaches 27 jobs at a susceptibility of 84.

Reach and exposure point the same way

04 / The concentration

So a hundred activities carry most of the field

The 100 most widely used activities account for 38 per cent of every job-to-activity link in the corpus. Automating that set would touch most of STEM work at once — which is where the leverage sits, and where a single scoring error propagates furthest.

100 activities · 38% of all links

Interlude / Side by side

Two jobs. What do they actually share?

Occupations overlap through the activity vocabulary, not through their task statements. Put two side by side and the shared spine separates from the speciality — and you can see whether what they have in common is the exposed part or the protected one.

03 / Two questions

Whether a machine can do the work is not whether it will.

The seven rated dimensions collapse into two that behave independently. Treating them as one number is the most common way this analysis goes wrong.

Fig. 03 — Exposure against anchoring 01 / 04
Exposure can a machine do it
Positions are relative to other STEM occupations.
01 / Capability

Exposure asks what a machine could do

How much of the cognitive core of this work could a current model perform, given the right inputs and tools? Document handling, retrieval, drafting and analysis score high. Work requiring hands in the world scores low.

Highest — proofread documents · exposure 95

02 / Permission

Anchoring asks whether a human must own it

Accountability, the cost and irreversibility of error, judgment under uncertainty, and whether the relationship is the substance of the work. A surgeon and a data analyst can both be highly capable targets and sit at opposite ends of this axis.

Highest — operate on patients · anchoring 80+

03 / Independence

The two barely correlate

Drafting a legal opinion is almost entirely language work and still requires a named human to answer for it. Ninety-two of the 963 activities score high on both. A single “automation risk” number cannot represent them.

Exposure vs anchoring — r = −0.03

04 / The split

The field divides almost in half

Ninety-eight STEM occupations sit in the displaceable quadrant and a hundred in the human-anchored one. Which half a job lands in is decided more by accountability than by capability.

98 displaceable · 100 human-anchored

04 / The frontier

Capability is not what is holding most of this work.

Watson's handoff framework scores work on tractability — whether AI can lead it — and resistance — whether it will be permitted to. The interesting cases are where those two disagree.

Fig. 04 — Tractability against resistance 01 / 04
Tractability can AI lead it
Framework: Watson 2026. Curve calibrated to this corpus.
01 / Two axes

Can it, and will it be allowed to

Recurrence, feedback, observability and structure decide whether AI can lead a decision. Stakes and legitimacy decide whether it is permitted to. The first four are engineering questions; the last two are not.

Six of Watson's eight properties are covered here

02 / The line

The frontier is where work crosses

Below the curve, the handoff has happened. Above it, a human still leads. Highly tractable work crosses even at some consequence; work that is not tractable stays human-led however low the stakes.

36 handed off · 98 crossing now

03 / Watch points

Capability present, accountability holding

Twenty-nine occupations sit above the frontier with a wide gap between what AI could do and what is deployed. Watson reads that gap as willingness to permit the handoff — and an actor with looser norms can close it first.

29 watch points

04 / Who they are

Mostly clinical judgment and risk

Genetic counsellors, actuaries, preventive medicine physicians, epidemiologists. In each the analytical core is well within reach and a named human is still required to sign.

Mean willingness gap among them — +45

05 / The ladder

Eight occupations have crossed. A hundred and forty could.

Watson scores cognitive leadership on six stages, from human-only to AI-led and unreviewed. A handoff is a crossing between stages. Scoring what is deployed and what is reachable separately shows how much travel is pending.

Fig. 05 — Stage today, stage reachable 01 / 04
Six stages of cognitive leadership
Stage is the lower of the capability and permission ceilings.
01 / The scale

From informed to unreviewed

Human only. AI informed. AI recommended. AI executed with a human veto. AI led with a human audit. AI led, unreviewed. Each step moves authority, not just capability.

Six stages · a handoff is a crossing

02 / Today

Most STEM work sits at recommendation

On what is actually deployed, 158 occupations sit at “AI recommended” and 101 at “AI informed”. Eight have reached “AI executed with a human veto”.

8 occupations at AI-executed today

03 / Reachable

Capability has already moved past that

Rate the same occupations on what a current model could do rather than what is deployed, and 140 reach “AI executed with a human veto”. The distance between those two numbers is pending handoff.

140 could be there now

04 / The quiet one

The second crossing has no event

Watson expects two crossings to carry most of the weight: proposing action to taking it, and human veto to after-the-fact audit. The second happens by erosion — review thinning toward ceremony, with nothing to observe.

120 occupations sit at one of the two

06 / The people

Twenty-one million workers, pulling in opposite directions.

Occupation counts weight a nurse the same as an actuary. Joining BLS employment asks a different question: not which jobs are exposed, but how many people are in them.

Fig. 06 — Workers by handoff quadrant 01 / 04
Workers covered 21.5 million
BLS OEWS national file. Employment counted once per SOC code.
01 / The count

21.5 million across 195 occupation codes

O*NET reports at a finer grain than BLS, so several O*NET occupations roll into one SOC code. Attaching employment to each would count 3.4 million registered nurses five times over. Susceptibility is averaged up first.

195 SOC codes · 37 of them collapsed

02 / The split

Forty-three per cent in the displaceable half

Weighted by headcount, 9.2 million STEM workers sit in the displaceable quadrant and 8.9 million in the human-anchored one — a wage bill of 1.1 trillion dollars on the exposed side.

43% displaceable · 41% human-anchored

03 / The cancellation

The two largest occupations disagree

Registered nurses, 3.4 million people, sit at the human-anchored end. Software developers, 1.7 million, sit near the top of the exposed end. They very nearly cancel.

Nurses susceptibility 50 · developers 69

04 / The null result

Weighting barely moves the answer

Employment weighting shifts mean susceptibility by half a point. Headcount is not concentrated at either end, which means the unweighted ranking was not misleading. That is worth reporting precisely because it could have gone the other way.

Mean susceptibility 58.2 → 58.7

08 / The premium

What protects well-paid work is not that AI cannot do it.

The Frey and Osborne era found automation risk falling as wages rose. Run the same test against LLM exposure and the two components of susceptibility pull in opposite directions.

Fig. 08 — Exposure and anchoring by wage decile 01 / 04
Wage vs exposure r = 0.03
Each decile holds 2.15M workers; r values are across occupations.
01 / Capability

Exposure does not care what a job pays

Across the 195 occupation codes with wage data, the correlation between pay and exposure is 0.03. Whatever decides how much of a job a model could do, it is not the salary.

Wage vs exposure — r = 0.03

02 / Permission

Anchoring is what tracks pay instead

Across those same occupations, accountability and the cost of error correlate with the wage at 0.40. Better-paid STEM work is not harder for a model to attempt — it is work someone has to answer for. The line on screen is employment-weighted and does not climb smoothly: it dips through the software and engineering-management deciles and rises again at the clinical top.

Wage vs anchoring — r = 0.40 across occupations

03 / The peak

Exposure peaks just below the top of the scale

Susceptibility climbs from 52 in the bottom wage decile to 69 in the seventh — and then falls back to 58 at the very top. The highest-paid decile is physicians, dentists and specialists, where anchoring climbs again. The most exposed workers are not the best-paid; they are the well-paid tier just beneath them.

Peak at the 7th decile · 69, falling to 58 at the top

04 / The share

And accountability protects a minority

Sorting every occupation by what is actually holding it: 42 per cent of STEM workers are in work a model largely cannot do, and only 9 per cent in work it could do but is not permitted to. The accountability premium is real, and it is narrow.

9% of workers protected by accountability

09 / Nowhere adjacent

The jobs next door are exposed too.

The standard answer to displacement is to move into an adjacent occupation. That assumes adjacency and exposure are independent. In the activity network they are not.

Fig. 09 — Where an exposed job could move 01 / 04
Moves that exist 107 of 268
Destination must share activities and be meaningfully less exposed.
01 / The test

A destination has to clear three bars

Enough shared activities that the move is plausible. Meaningfully lower susceptibility, not noise. And the shared activities have to include the destination’s protected work — otherwise the worker carries their exposure with them.

Overlap · relief · direction

02 / The moves

Where a move exists, it is usually genuine

A hundred and seven occupations have a destination that clears all three. Ninety-eight of those are real moves into better-protected work; nine share only the exposed half, and would carry the problem along.

98 real · 9 carry the exposure

03 / The stranded

But most occupations have nowhere to go

A hundred and sixty-one of 268 have no close neighbour that is meaningfully safer. Fifty-nine of those are themselves highly exposed, covering 11.5 million workers. Database administrators, data scientists, programmers and web developers sit in a neighbourhood where everything is exposed.

11.5M workers exposed and stranded

04 / How firm

The direction holds; the number does not

Sweeping the thresholds moves the stranded count between 105 and 240 of 268. How much relief you demand changes the answer a great deal, and how much overlap you require barely changes it at all. The finding is that exposure is clustered — not that the number is 161.

Sweep — 105 to 240 stranded

10 / Has it moved yet

The work has already started changing — where anyone looked.

Every other measure here rests on a model’s judgment about what could happen. This one does not. O*NET archives every release, so the task statements attached to a job can be diffed across eleven years and the turnover counted directly.

Fig. 10 — Task turnover, 2015 to 2026 01 / 04
Turnover since 2015 5.2% overall
175 occupations tracked; 93 skipped on the 2019 SOC revision.
01 / The raw number

Almost nothing changed

Across 175 STEM occupations present in both the 2015 and 2026 releases, 3,553 task statements became 3,666. Two hundred and forty-three were added, 130 retired, and 3,423 survived untouched. Seventy-four occupations have identical task lists to eleven years ago.

5.2% turnover · 6.6% of today’s tasks are new

02 / And no inflection

The turnover did not accelerate after 2022

If language models had already reshaped these jobs, the diff would show it. The largest single step in the series is 2019 to 2021 — before ChatGPT. Every step since 2022 is smaller than the ones before it.

Largest step 2019–2021, at 2.5%

03 / The confound

But O*NET only re-surveys on a rolling cycle

An occupation whose tasks did not change may simply not have been looked at. Splitting on the date O*NET last reviewed each one: the 153 re-surveyed since 2022 turned over 8.1 per cent, the 115 that were not turned over 2.4. The headline figure is diluted by occupations nobody checked, and no amount of care with the diff fixes that.

8.1% where checked · 2.4% where not

04 / The signal

And exposed jobs moved roughly three times faster

Among occupations O*NET did re-survey, the highly exposed ones turned over 13.2 per cent of their tasks against 4.6 for the least exposed. Small sample — sixteen occupations — and O*NET does not pick what to re-survey at random. But it points the same way the scores do, from data that knows nothing about them.

13.2% exposed · 4.6% not · n=16

11 / One job, task by task

What happens to each task, and what is left.

The same bundle, coloured by what becomes of each task under a chosen set of assumptions. Nothing here is a forecast — the scenario sets how much accountability an actor is willing to hand over, and everything else follows from the scores already measured.

Fig. 11 — A single job’s tasks by fate 01 / 06
A bundle of tasks 21 tasks
New tasks are observed from O*NET, not modelled.
01 / The bundle

Start with everything the job contains

O*NET lists each occupation as a set of task statements. Laid out flat they are just a bundle — no ordering, no weighting, and nothing yet said about which of them a machine could take.

Every square is one documented task

02 / What stays

Some of it a model cannot touch

Work needing hands on a patient, presence in a room, or a judgement someone has to own. These stay grey at every setting — no scenario moves them, because the constraint is not capability.

Grey at every scenario

03 / What gets help

Some gets faster rather than taken

A model drafts, summarises, plans or checks, and a person keeps the work. This is the largest category in the middle scenario — the least dramatic result in the dataset, and the most plausible.

The largest group under Substantial

04 / What gets taken

And some of it goes

High exposure, little accountability attached: retrieval, formatting, routine documentation, scheduling. Move the scenario and watch this category eat into the middle one. That movement is the whole argument.

21% → 45% → 69% of all tasks

05 / What arrives

New work appears too

Not modelled — observed. O*NET flags newly emerging task statements, 121 of them across these occupations. For a nurse midwife: evaluating patients’ mental health, screening for gynaecologic conditions.

121 new tasks recorded in O*NET

06 / What is left

The job is not smaller, it is different

What remains gets more room. The tasks a model touched carry more volume per hour of human attention; the ones it cannot touch are what the job becomes. Whether that is a better job is not a question this data can answer.

Scale shows capacity, not headcount

12 / Where people end up

Reshaped is not the same as displaced.

Weighted by employment, and crossed with whether a worker’s occupation has anywhere adjacent to go. Move the scenario and both halves move — but not by the same amount, which is the point.

Fig. 12 — Workers by outcome 01 / 04
Workers reshaped 43.8%
Employment counted once per SOC code. Undated.
01 / Reshaped

Start with whose job changes at all

An occupation counts as reshaped when half its task list or more falls into the automated category. Under the middle scenario that is 35 per cent of STEM workers; under the modest one 12; under the extreme one 83.

12%44%83% of workers

02 / Little change

Most of the workforce, at most settings

The remainder are in occupations whose documented task list survives the threshold largely intact. That is the majority everywhere except the extreme case, and it is the part headlines tend to drop.

The majority in two of three scenarios

03 / Could move

Some have somewhere to go

Crossing the reshaped group with the transition map: these workers are in occupations with a close, meaningfully less exposed neighbour that shares the destination’s protected work.

2% → 12% → 44% of workers

04 / And some do not

The rest are in a neighbourhood that is all exposed

This is the number that does not scale kindly. Between 9 and 34 per cent of STEM workers are in reshaped occupations with no adjacent destination. Under the modest scenario, four in five of those affected have nowhere to go.

9%32%39% stranded

07 / The check

A model rating work is an assertion until someone checks it.

Every other number here is internally consistent by construction. This is the only part that could have come out wrong in a way the rest of the pipeline would not catch.

Fig. 07 — Against human expert ratings 01 / 03
A model rating work is an assertion
Benchmark: Eloundou, Manning, Mishkin & Rock (2023).
01 / The problem

Self-consistency is not evidence

The scores order the model's own categorical verdicts monotonically, which is reassuring and proves nothing — both come from the same model. An external benchmark is the only thing that can fail.

Convergent validity 74.9 > 65.9 > 43.6 > 40.3

02 / The benchmark

Human annotators rated the same occupations

Eloundou and colleagues published exposure ratings keyed to the same O*NET codes, from human annotators as well as a model. All 268 of our occupations matched with no crosswalk.

268 of 268 matched

03 / The result

They agree closely

Correlation of 0.85 with the human ratings. The weaker agreement with their no-tools definition is expected: this rubric explicitly asks what a model could do given the right tools.

r = 0.85 against human experts

What this cannot tell you

Three limits worth stating before anyone cites it.

The unit is the task, not the decision. Watson's framework scores recurring decisions, and he is explicit that identifying which tasks are decisions is the novel work. It has not been done here, so the stage numbers are provisional.

The scores are model judgment. They agree with human annotators at r = 0.85, which is evidence of validity, not a substitute for it. No reliability estimate exists yet — nobody has checked whether a second run agrees with the first.

Exposure is not displacement. Nothing here measures what employers will do, what regulation will permit, or how fast anything diffuses. It measures the shape of the work and where accountability sits.