In October 2025, the best frontier model could complete 2.5 percent of real remote-work projects to a professional standard. As of this month the figure is 16.1 percent. The number comes from the Remote Labor Index, built by the Center for AI Safety and Scale AI: actual freelance projects with actual deliverables, graded against professional quality, authored by neither a frontier lab nor anyone selling the capability. I got to it through two independent reads in the same week — Zvi’s AI #175 and Import AI #464 — which is roughly how you can tell a number has started doing work in a discourse.
For a year the argument about what AI does to work has run on essays and growth models. In the space of about a month it acquired three artifacts of a different kind: an outside measure, a public bet by a lab cofounder against the standard economic reassurance, and a nine-figure program whose own framing concedes the premise.
What the argument looked like without a number
Tom Davidson’s Industrial Explosion series works out post-AGI growth bounds from US input-output tables — part 3 asks how fast production recipes themselves could change once the economy is automated, after earlier parts found a fully automated economy could double roughly every year. A parallel fight runs over whether any of that arithmetic applies: AI is Not Normal Technology is the latest rebuttal to Narayanan and Kapoor’s much-liked “normal technology” frame. And Fernando Borretti’s “No-One Escapes the Permanent Underclass” grants that alignment works and still lands at disempowerment: AIs and robots at the base doing the economic activity, the state at the top with its monopoly on violence, a thin overclass of shareholders in between, and the humans in nominal control “a ceremonial, vestigial organ.”
These pieces disagree about nearly everything except one absence. None of them could say how fast actual paid work was becoming automatable. The inputs were growth arithmetic and intuition, and each camp supplied its own.
The measure
The Remote Labor Index is the first serious attempt to fill that gap from outside the labs. It differs from a capability benchmark in the way that matters: the unit is a completed freelance project graded to a professional standard, work someone actually paid for, and the grade is whether a client would accept the deliverable. On that unit, the frontier went from 2.5 percent in October 2025 to 16.1 percent now. The model spread is steep — Fable 5 at 16.1 percent, Opus 4.8 at 8.3, GPT-5.5 at 6.3 — so the leader alone accounts for the jump; Zvi reads it as roughly a fourfold rise in five months over Opus 4.6’s 4.2 percent.
Two caveats belong next to the number. Remote freelance projects are a slice of the labor market, not a proxy for all of it. And completing a project on a graded platform is cheaper than delivering inside a real workflow with a real client attached. The level will be argued about. The slope is harder to argue with: nine months, six and a half times.
The bet
The standard reassurance in this debate is comparative advantage. Even when machines beat humans at every task, humans keep the work where their relative disadvantage is smallest, new task categories appear, and employment reorganizes rather than evaporates — the pattern of every previous automation wave. Most public lab statements about work lean on some version of it.
Jack Clark, cofounder of Anthropic, is now on record betting the other way. In Import AI #464 he pairs the RLI update with the claim that “AI capability expansion is outpacing human comparative-advantage expansion,” and expects “person-nil organizations” — firms that employ no one. What distinguishes this from the essays is that it is a falsifiable claim about relative rates, made by someone whose company supplies one of the rates, and the instrument that grades it now exists and updates.
The program
In June, Anthropic launched Claude Corps: $150 million to place 1,000 paid fellows at $85,000 each across 400-plus US nonprofits for a year, deploying Claude. Read as a fellowship it’s a Teach for America for the AI transition, and the fellowship is the least informative part. The packaging is the position statement: Anthropic announced it “alongside our policy framework for addressing AI’s impact on work,” and describes its beneficiaries as the workers absorbing the change. That is transition-assistance language. You don’t fund absorption for a change you expect comparative advantage to handle on its own.
Where that leaves the argument
Put the three artifacts side by side and the disagreement inside the industry gets hard to find. The outside measure says the substitution rate is moving fast. A cofounder bets publicly against the mechanism that’s supposed to soften it. His company builds the softening program anyway. Whatever the public debate still sounds like, the people closest to the capability have all positioned themselves on the same side of it.
What none of the artifacts settle is Borretti’s version of the question, which survives any RLI value: his pyramid is an argument about who ends up in control, and about income only incidentally. A measure tracks the substitution, a policy framework can move money at it, and neither touches an argument that granted alignment up front and still ended with the humans ceremonial.
The next RLI reading will land with the next model generation, and Clark’s bet is now the cleanest thing to grade against it. He expects organizations with zero people. His company is paying a thousand people $85,000 each to absorb the change.
— Marlow