The Remote Labor Index maintainers claim GPT-6 Astra can now automate 20.8% of randomly sampled remote projects, up from 2.5% in last October's results โ an eightfold jump in measured remote-work automation within a year.
state: seedheat: mediumuncertainty: mediumconvergesscott: highremote-labor-index agent-evaluation gpt-6-astraOpenAI
What is this?
The Remote Labor Index (RLI) is a benchmark from the Center for AI Safety (CAIS) and Scale AI, released October 2025, that measures how often AI agents can complete real, paid freelance projects (3D/CAD, design, video, data analysis, web apps; 240 projects worth ~$144k) at a quality a paying client would accept, judged by human evaluators against gold-standard deliverables. Its baseline result was that the best agent (Manus) automated just 2.5% of projects. The new case is a CAIS announcement that OpenAI's GPT-6 Astra now automates 20.8% โ but the supplied snippets conflict internally: the safe.ai blog post cites a rise to 15.8%, The Decoder reports 16% in eight months, and only the CAIS social post states 20.8%, so the exact figure and which model it attaches to are not consistently established by the supplied material. Notably, the successful projects skew toward generative-from-scratch work (audio, logos, reports), with agents still failing multi-step briefs and precise editing.
Why it matters to Scott
The RLI result is exactly the evidence base the Terminal Value Doctrine predicts but has lacked: a third-party, paid-work benchmark (graded against client-acceptable deliverables, i.e. measuring the right unit per benchmarking-the-wrong-unit) showing measurable substitution of purchased remote labor with AI in a single year โ direct quantitative support for the Externalisation Share thesis and a dated-receipts publishing opportunity. It also lands on Scott's own history: RLI automates the oDesk/offshore freelance market he personally hired in, and the failure pattern (wins at generative-from-scratch work, losses on precise multi-step briefs) matches his one-shot vs. long-horizon distinction. Caveat for the writeup: the supplied material conflicts on the headline number (20.8% vs 15.8% vs 16%), and the radar does not yet track RLI itself โ this is a new story, not a repetition.
ip:framework.terminal-value-doctrine-for-professional-servicesip:concept.demand-side-disintermediationip:concept.benchmarking-the-wrong-unitip:framework.progressive-resolutionwork:project.odesk-comradar:concept.agent-benchmarksradar:concept.agent-evaluationradar:concept.agent-economics
queries asked of Scott's wikis
- agent evaluation real-world task completion vs synthetic benchmarks
- coding agent harness reliability long-horizon task failure modes
- benchmark contamination and metric gaming in agent evals
- AI freelance and remote labor market automation economics
- generative one-shot creation vs precise multi-step instruction following
- GDP-val and economic-value measures of AI capability
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 437h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p53 vs 1032 stories at the 336h mark (now 437h old) โ ahead of astra-skills-prompt-migration (1.1x), behind autobot-persistent-chatgpt-harness (0.9x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-23T17:19:24Z
grounded: converges/high โ The RLI result is exactly the evidence base the Terminal Value Doctrine predicts but has lacked: a third-party, paid-work benchmark (graded against client-accep
2026-09-23T17:13:01Z
case created โ A first-party benchmark claim of an order-of-magnitude automation jump is a major, checkable capability assertion that no open case yet tracks.
Decision trace
- 09-24 03:19groundThe RLI result is exactly the evidence base the Terminal Value Doctrine predicts but has lacked: a third-party, paid-work benchmark (graded against client-acceptable deliverables, i.e. measuring the r
- 09-24 03:13createA first-party benchmark claim of an order-of-magnitude automation jump is a major, checkable capability assertion that no open case yet tracks.