2026-10-11 18:01 UTC

NeoCognition's ApprenticeBench claims to measure whether AI agents can learn and perform a real job end to end; adoption by evaluators would establish it as a reference benchmark for long-horizon job competence.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation long-horizon-agentsNeoCognition

What is this?

ApprenticeBench is a benchmark released in September 2026 by NeoCognition, a research org, that evaluates whether AI agents can learn a real knowledge job end to end: the agent plays an accounts-payable clerk at a simulated California construction company (Acme Home Builders), processing 100 vendor bills in the Odoo ERP with only a handbook, six months of historical records, and mentor feedback โ€” onboarding feedback that tapers to sparse month-end reviews. It claims to be the first benchmark combining computer use, continual learning, and long-horizon agency in a realistic job, and its blog reports a capability jump with Fable 5.1 and GPT-6 Astra, which reportedly beat the best human tester by ~20%. The adoption claim in the hypothesis is harder to verify from these snippets: coverage beyond the leaderboard aggregator (BenchLM) is thin, and most of the substance echoes NeoCognition's own pages; a competing open-source agent benchmark (Brackett's Agent Effectiveness Index) launched the same week, so 'reference benchmark' status is not yet established by this material.

Why it matters to Scott

ApprenticeBench independently builds what Scott's own benchmark practice argues for: an evaluation of real-deployment job competence (continual learning + computer use + long-horizon work in a live ERP) rather than narrow static tasks โ€” the same scored-episodes-into-memory pattern as his callsimulator/self-play work and his trace-backed, own-fixture benchmarking, now done at job scale by an outside org. It matters because if it gains the reference status the hypothesis claims, it becomes an external yardstick he could run his harness designs against and strengthens his benchmarks-must-measure-real-work position; but the radar already tracks a crowded field of near-cousins (Real-SWE's enterprise-context benchmark, Wharton-Harvard business benchmark, LongHorizon-Harness), and adoption is unverified in the snippets, so this is a credible convergent data point rather than a story-changing event.
dev:concept.llm-self-play-refinementdev:project.callsimulatordev:concept.trace-backed-agent-comparisondev:project.remote-execradar:concept.agent-benchmarksradar:concept.long-horizon-agentsradar:concept.continual-learningradar:concept.computer-use-agentsradar:specific-real-swe-release
queries asked of Scott's wikis
  • agent memory: memory notes, continual learning, learning from experience across tasks
  • agent harness design: long-horizon task loops, feedback signals, error correction
  • benchmark criticism: what benchmarks measure vs real deployment, Goodhart
  • agent evaluation tooling: how do I test whether my agents improve over time
  • computer-use agents: ERP / GUI automation, browser and desktop control
  • on-the-job learning vs RAG: procedural knowledge acquisition, experience replay

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady1 platformsage 430h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-23 17:47โญ origin directly observedApprenticeBench: Can AI agents learn a real job end to end?
moyicat on hacker news
โ€”
09-23 17:47amplified on hacker news ๐Ÿ‘‘hn.story.49819853
moyicat
peak 2 ยท 1 comments ยท 101% of case engagement
09-23 21:22our radar first saw it ยท +3.6hdiscovery anchor: hn.story.49819853โ€”
pace: p32 vs 1032 stories at the 336h mark (now 430h old) โ€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hn โญApprenticeBench: Can AI agents learn a real job end to end?moyicat21

Interpretation history

Decision trace