So far the argument about AI and jobs has been largely fought with survey data, company announcements and anecdote. But this month a US federal statistical agency put administrative numbers on the table, and the picture they show is not ambiguous.

What happened to the graduates
Three economists at the US Census Bureau (Cody Orr, Lee Tucker and Lawrence Warren) published a working paper called “Graduating into Disruption”. They took administrative records covering roughly 29 per cent of all US bachelor’s degrees awarded between 2016 and 2024 (about 6.7 million graduates) sorted their majors by how exposed those fields are to AI, and watched what happened after ChatGPT arrived at the end of 2022.
Graduates from the most exposed tenth of majors saw their chance of landing an initial job fall by five percentage points, and their first full quarter of earnings fall by thirteen per cent, compared with graduates from less exposed fields. The authors’ own benchmark for that number says:
“Comparable in magnitude to the earnings losses
associated with graduating into a large recession”
A recession nobody declared, but for just one single group of people.
Roughly half of that earnings decline is not people being paid less for the same kind of work. It’s people ending up in different work altogether - a shift into lower-paying sectors, and the paper names restaurants and retail specifically. The other half is lower pay within the same industries. So the effect is not simply “graduates earn a bit less”. A substantial share of it is graduates who expected one kind of career and are now doing something else. This is exactly how these people get hidden the reports that claim an “AI jobs boom”.
The effects fade the further people get from labour-market entry, but stay substantial for the most exposed fields. Erik Brynjolfsson, whose own research team has been tracking a similar pattern, expects the effects to be “quite a bit more noticeable” by the next graduating class.
It’s worth noting that this is a working paper that has not been through the review that Census Bureau publications normally get, and the Bureau says so explicitly. The measurement is relative (most-exposed majors against less-exposed ones) not an absolute count of jobs lost. And the timing of ChatGPT’s launch sits awkwardly close to the 2022-23 interest-rate rises and the tech layoffs that followed, which hit computer science and business graduates for reasons that had nothing to do with AI. The authors handle that partly by comparing exposure groups against each other, which nets out an economy-wide downturn, and by showing the trends were flat before 2020 and diverged immediately at the end of 2022.
What makes this significant is that this is the fourth independent method we’ve tracked that finds this pattern, and it’s also the first to come from a government statistical agency working with administrative records rather than surveys.
Why it hits the entry level specifically
Think about why junior roles exist in professional work. A graduate lawyer does not get handed the case strategy task. They get the document review, the first-draft memo, and the precedent search. A junior analyst gets the data audit, not the investment decision. A first-year accountant gets the reconciliation, not the tax advice. In every case, the newcomer is given work whose output can be checked by somebody senior.
That’s the organising principle of professional training. Give the newcomer the tasks where mistakes are catchable, check the output, and then gradually extend this trust. The apprenticeship structure of skilled work is, underneath it all, a verification system.
And that is exactly where an “AI with verification” system fits in.
That suggests the question of which jobs go first is less about how “cognitive” the work is, and more about how cheap it is to check the answer. Where verification is easy (does the code pass the tests, does the number reconcile, does the extracted field match the source document) a machine that is unreliable but free can be run repeatedly until something passes. Where verification is expensive (judgement, strategy, persuasion, relationships, knowing which question to ask, being accountable when it goes wrong) that trick stops working.
It also explains something the older “routine versus non-routine” framing never quite did. Junior professional work is not routine. It is difficult, varied and demanding. But it is checkable, and I think that turns out to be the property that really matters.
All of that happened in the tier we can see
Everything above was measured against a version of AI that has a price tag.
It cost money per seat. It needed a procurement decision, a vendor relationship, a budget line and somebody’s approval. Or at the very least a subscription. It ran in somebody else’s data centre. It was metered, logged, and realistically available mostly to employers large enough to buy it, or individuals wealthy enough to pay for it.
That is the tier we can count. But now, there is a second one.
The split
In September, a company called PrismML released Ternary Bonsai 2, a compressed version of Alibaba’s Qwen3.8-27B open-weight model. The original is a 27-billion-parameter model that needs about 54 gigabytes to run. The compressed version is just 5.95 gigabytes. It is released under the Apache 2.0 licence, and it has been downloaded over 2 million times so far. And most importantly, it runs at good-enough speed on a current laptop.
Two things make this more than a single product release.
The first is that the conversion has become cheap and routine. The technique used to be something you had to train a model for from scratch, and the results were small and weak. During 2026 the field worked out how to compress existing strong models after the fact. An Intel method presented at ICML this year does it using around 500 calibration examples - which the authors describe as roughly a hundred-thousand-fold reduction in the training data required. Research from Meta explains why this particular level of compression works better than the more obvious alternatives. The practical consequence is that now, every strong open-weight model released from here on could become laptop-class within weeks, cheaply, by anyone who wants to do it.
The second is that you do not need the model to get better in order to get better answers. You can run a free model many times over and keep the best result - which converts spare computing time into quality, and spare computing time on hardware you already own is close to free - or more accurately zero marginal cost. And the scaffolding around a model now matters enormously just on its own. When OpenAI’s GPT-6 Astra was tested on the ARC-AGI-3 reasoning benchmark this month, the model by itself scored 62.7 per cent, while the same model with tooling around it scored close to 99. Same model. Different harness.
Put those together and you get capability that can keep rising without any new model release, without a per-seat cost, on hardware that has already been bought.
But this is not the cheap version catching the expensive one.
Mozilla’s State of Open Source AI, updated in September, takes METR’s data on how long a task a model can handle and fits separate trend lines to open and closed models. The distance between those two lines is not a score - instead it is a calendar. Mozilla’s figure is about 4.4 months. Their headline puts it plainly:
Closed models handle 8-to-12-hour tasks,
and open models get there four months later.
A fixed four-month lag produces a growing score gap whenever the frontier curve steepens - which is exactly what has been happening. The lead is real, but it is not compounding. Mozilla’s phrasing is that the gap “resets every release cycle” rather than accumulating. My view is that it breathes: it opens when a frontier lab ships something, and closes again over the next few months.
The consequence is the important part:
“The band moves, because a task at the closed frontier today
is at the open frontier four months later” - Mozilla
So the frontier keeps its lead, but the capability needed for any particular job gets commoditised over the following months. That is a very different thing from a moat.
The three layers
The split is not cheap-versus-expensive. It is better seen as two classes of thing, and one of them can be deployed in two different ways.
Frontier cognition is the top tier: expensive, metered, logged, guardrailed, centralised, still accelerating. On the Artificial Analysis Intelligence Index in early September, the top four models were all closed - and the next four were all open-weight.
Commodity cognition is the second class: smaller models, good-enough for a great deal of real work, running on ordinary hardware. It is not trying to be the frontier. But it arrives at the frontier’s position a few months later and at just a fraction of the cost.
And commodity cognition shows up in two very different places:
Submarine cognition runs on hardware somebody already owns: A laptop, a workstation, a server in a cupboard. Free at the margin, unmetered, unlogged, invisible.
Utility cognition is the same class of model running on somebody’s cloud: Cheap, but billed, logged and countable.
That distinction is critical. Cheap does not mean invisible. Utility cognition is cheap and countable. But the submarine tier is genuinely out of sight.
The thing that ties it together: routing
Here is the mechanism that makes all of this economically consequential.
The harness around a model does not only improve quality. It allocates. Break a workflow into steps, send each step to the cheapest layer that can do it acceptably, and keep the expensive layer for the parts that genuinely need it. Today’s agentic systems can already do this automatically, but mostly by choosing between frontier models. But as the commodity layer becomes good-enough, the growing amount of this work will likely route downwards.
Which means the commodity layer does not need to match the frontier for the economics to break open. It only needs to be good-enough for the share of work that gets routed to it - and increasingly the harness works out that share by checking the results.
This closes the loop with the graduates. You can only route work down to a cheaper layer if you can tell whether the cheap answer is good-enough. Cheap-to-check work routes down. Expensive-to-check work stays at the frontier, or stays human.
Verifiability decides what gets routed. Routing decides what it costs. Cost decides how far it reaches.
And the same property that makes junior professional work substitutable (that somebody was already checking it) is the property that makes it routable.
Why the split matters more than any one layer
A moving lead, not a moat. The frontier stays ahead, but the capability level needed to do a particular job keeps arriving at the commodity layer a few months later. “The frontier is still ahead” only reassures you about the slice of work that needs the frontier. Whil the frontier’s long tail is being eaten.
Different reach. The frontier tier needs a purchase decision. But the commodity tier only needs a laptop somebody already owns, or a cloud bill small enough not to need approval. For the enormous long tail of small and medium employers, the binding constraint was never the price per query - it was having to decide to buy something at all, or the fear of token shock. Routing can remove what remained of the cost argument.
Different visibility. We can count two of these three layers. Frontier and utility both produce invoices, telemetry and vendor records. Submarine produces none. When my research turned up how thin the public data is, the result was starker than I expected: a full-text search of the entire Stanford AI Index 2026 (the field’s flagship annual report, some 34,000 lines) returns zero mentions of on-premises inference, local inference, on-device inference, or any of the main local-model tools. Of course this is somewhat obvious because we are only really just passing this threshold now. But the key point is that this hole in the reporting data is unlikely to change. Nobody publishes a split between AI that runs in a data centre and AI that runs on someone’s desk. The measurement just does not exist and it is likely to remain unmeasurable.
Different regulatory exposure. Every serious proposal on the table (the pacing regime the lab chief executives asked for earlier this month, licensing, evaluation requirements, compute thresholds) works by metering. That reaches frontier and utility. It does not reach a model sitting on a laptop with no vendor in the loop. The “thing being regulated” and a core part of the “thing spreading” are coming apart. Meanwhile nations at the United Nations General Assembly are calling for control.
Different work. These layers do not compete for the same jobs so much as divide them. The frontier takes the hard, long-horizon, high-stakes work inside firms that can pay for it. The commodity layers take the enormous volume of checkable work, everywhere, at very little cost. They are not substitutes for each other. They compound against the same labour market from different directions, and the harness is likely to decide which direction each task comes from.
What we do not know
The quality claim is contested. PrismML says its compressed model retains 98.2 per cent of the original’s ability. That figure is self-reported and I could not find any independent replication. Meanwhile an independent academic team compressing a smaller model the same way reports retention of about 78.5 per cent - a loss of nearly eight points. The distance between those two numbers is the distance between “replaces the work” and “helps with the work”. PrismML’s own smaller model concedes a larger deficit than its flagship claims, which is worth noting.
The thing that really matters is untested. Every one of these benchmark figures measures single-turn question answering. Substituting for real work needs something different: multi-step tool use, holding coherence over long tasks, output that does not drift halfway through a workflow. Compression damage is known to concentrate in exactly those long-horizon behaviours. Nobody has rigorously run that test publicly - yet. Until somebody does, the strongest version of this argument is unproven.
What I am confident about: capability that was frontier-class a year ago now runs on a laptop, the ceiling of what fits in six gigabytes is rising.
But currently, measuring the uptake is harder than it should be. The best I could manage was adding up the per-model download counters on Ollama’s library myself - which gives a little over a billion cumulative pulls, against about 420 million a year earlier. Hugging Face points the same way from a different angle with the top 1,200 repositories in the file format these local tools use drawing more than 240 million downloads in a single month. Both only count downloads, not people, and neither tells you how much is real work rather than curiosity. That I had to construct these numbers myself is partly the point right now.
Obviously, this has not yet directly impacted employment because we are just passing this real threshold now. But it is a real shift in the landscape that deserves to be named and tracked because it is likely to be very significant.
Where that leaves things
The graduates in the Census study faced a version of AI that had a price, a vendor and a procurement process. Only employers large enough to buy it could use it against them. But that version cost them something close to a recession.
The next cohort meets a different arrangement. A frontier layer that keeps pulling ahead for anyone who can afford it. And a commodity layer that arrives at the frontier’s position a few months later for a fraction of the price, running either on someone’s cloud or on hardware already sitting on the desk - and the agile ones will apply it themselves. Meanwhile, software in the middle that is getting steadily better at deciding which layer each piece of work should go to.
That last part is the bit worth watching. Not whether the cheap models catch the expensive ones (they do not need to) but how good we get at spraying the work across the layers.
And for one of those layers, we currently have no measurment instrument at all.

