2026-10-11 16:38 UTC

The Remote Labor Index maintainers claim GPT-6 Astra can now automate 20.8% of randomly sampled remote projects, up from 2.5% in last October's results โ€” an eightfold jump in measured remote-work automation within a year.

state: seedheat: mediumuncertainty: mediumconvergesscott: highremote-labor-index agent-evaluation gpt-6-astraOpenAI

What is this?

The Remote Labor Index (RLI) is a benchmark from the Center for AI Safety (CAIS) and Scale AI, released October 2025, that measures how often AI agents can complete real, paid freelance projects (3D/CAD, design, video, data analysis, web apps; 240 projects worth ~$144k) at a quality a paying client would accept, judged by human evaluators against gold-standard deliverables. Its baseline result was that the best agent (Manus) automated just 2.5% of projects. The new case is a CAIS announcement that OpenAI's GPT-6 Astra now automates 20.8% โ€” but the supplied snippets conflict internally: the safe.ai blog post cites a rise to 15.8%, The Decoder reports 16% in eight months, and only the CAIS social post states 20.8%, so the exact figure and which model it attaches to are not consistently established by the supplied material. Notably, the successful projects skew toward generative-from-scratch work (audio, logos, reports), with agents still failing multi-step briefs and precise editing.

Why it matters to Scott

The RLI result is exactly the evidence base the Terminal Value Doctrine predicts but has lacked: a third-party, paid-work benchmark (graded against client-acceptable deliverables, i.e. measuring the right unit per benchmarking-the-wrong-unit) showing measurable substitution of purchased remote labor with AI in a single year โ€” direct quantitative support for the Externalisation Share thesis and a dated-receipts publishing opportunity. It also lands on Scott's own history: RLI automates the oDesk/offshore freelance market he personally hired in, and the failure pattern (wins at generative-from-scratch work, losses on precise multi-step briefs) matches his one-shot vs. long-horizon distinction. Caveat for the writeup: the supplied material conflicts on the headline number (20.8% vs 15.8% vs 16%), and the radar does not yet track RLI itself โ€” this is a new story, not a repetition.
ip:framework.terminal-value-doctrine-for-professional-servicesip:concept.demand-side-disintermediationip:concept.benchmarking-the-wrong-unitip:framework.progressive-resolutionwork:project.odesk-comradar:concept.agent-benchmarksradar:concept.agent-evaluationradar:concept.agent-economics
queries asked of Scott's wikis
  • agent evaluation real-world task completion vs synthetic benchmarks
  • coding agent harness reliability long-horizon task failure modes
  • benchmark contamination and metric gaming in agent evals
  • AI freelance and remote labor market automation economics
  • generative one-shot creation vs precise multi-step instruction following
  • GDP-val and economic-value measures of AI capability

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 437h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-23 10:58โญ origin directly observedLast October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%.
Puzzleheaded-King584 on r/OpenAI
โ€”
09-23 10:58amplified on r/OpenAI ๐Ÿ‘‘reddit.post.1wo2qw8
Puzzleheaded-King584
peak 21 ยท 2 comments ยท 100% of case engagement
09-23 11:20our radar first saw it ยท +0.4hdiscovery anchor: reddit.post.1wo2qw8โ€”
pace: p53 vs 1032 stories at the 336h mark (now 437h old) โ€” ahead of astra-skills-prompt-migration (1.1x), behind autobot-persistent-chatgpt-harness (0.9x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญLast October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%.
OpenAI
Puzzleheaded-King584212

Interpretation history

Decision trace