Anthropic — the San Francisco frontier-AI lab behind Claude — published 'Measurements for understanding the pace of AI development inside frontier labs,' introducing a prototype R&D Automation Index on Epoch AI's automation-level scale: in August 2026 Claude 'led' (AL4, end-to-end task execution under human supervision) 26% of Anthropic's internal AI R&D, up from under 1% in February, with >90% of measured work at 'collaborates' (AL3) or above and none fully autonomous; the ratings are self-evaluations using Anthropic's own models, Epoch supplied only the scale, and planned third-party evaluators are not yet operating. The supplied web snippets do not independently re-report that publication; they corroborate only the surrounding pattern via Anthropic's August 2026 Risk Report (Benzinga/Sahm Capital): Claude writes a 'large majority' of code merged into production codebases and internal research is faster with AI assistance — while Anthropic explicitly declines to claim a 2x speedup and admits its task-based evaluations are 'saturating.' Product-side moves in the same window (Claude Code defaulting to autonomous auto mode, Routines for cloud-scheduled agents, Managed Agents) and rival launches (OpenAI ChatGPT Work Agent, Google Jules, GitHub's coding agent) show the industry converging on supervised autonomous execution. The oversight figures (~30,000 concurrent internal agents, ~0.002% of >1B monthly decisions blocked, monitor flags escalated to humans) all trace to Anthropic's single self-report with no external validation, and the coverage spike peaked around 2026-09-18/19 before decaying to restatements.
| source | object | author | score | comments |
| 🟠 reddit | Anthropic reveals Claude is now leading 26% of its own R&D work, up from nearly zero 6 months ago singularity | Outside-Iron-8242 | 495 | 63 |
| 🟧 hn | Measurements for understanding the pace of AI development inside frontier labsRetrieved article excerptOpen article · Retrieved 2026-09-17T21:22:06.829281+00:00 Measurements for
understanding the pace of AI development inside frontier labs
AI systems are becoming exponentially more powerful and have begun to [automate more](https://www.anthropic.com/institute/recursive-self-improvement) of the process of building themselves. As the world considers [slowing the pace of frontier AI development](https://darioamodei.com/post/we-must-pace-the-frontier), the public needs more information.
In this post, we lay out measurement tools that can illuminate three critical aspects of AI development:
1. [The extent to which AI is building the next version of itself, as opposed to being built by humans](https://www.anthropic.com/institute/measuring-pace-of-ai-development#1-measuring-ai-led-ai-rd)
2. [Our ability to oversee and intervene in actions that AI agents take on Anthropic’s systems](https://www.anthropic.com/institute/measuring-pace-of-ai-development#2-measuring-oversight-of-ai-agents)
3. [The resources that power the development of more capable models](https://www.anthropic.com/institute/measuring-pace-of-ai-development#3-measuring-compute-allocation)
We also provide a snapshot of these metrics from inside Anthropic. It’s important to note that we would expect these numbers to shift if there were coordination on pacing the frontier, as called for by Anthropic CEO Dario Amodei. We plan to embed independent third-party evaluators from multiple organizations at Anthropic, and give them access to internal processes, systems, and data comparable to what internal risk assessment teams have. These third parties will verify safety practices, report incidents, and monitor key metrics such as the ones in this piece.
We are reporting these measurements because they give the public, third parties, and governments better visibility into the pace of AI development inside frontier labs. For each measurement, we describe what we measured, what the measurement showed, and what it would take to publish these measurements regularly in a form others can verify. We share methodological details in the Appendix.
## Reasons to track these measurements
The measurements in this piece are focused on *how models are built.* By better understanding the production process of models, we have a better chance of correlating model inputs, like compute, with model outputs, like capabilities. They complement capability evaluations, which measure *what models can do*. We publish those separately through our [Responsible Scaling Policy](https://www.anthropic.com/responsible-scaling-policy) (RSP) risk reports, which include evidence on how much our models are accelerating AI R&D. In our policy proposal on advanced AI, the [Advanced AI Framework (AAIF)](https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf), we propose rules of the road for how any lab releases safe models, including transparency obligations that governments could require, such as risk reports. Together, these proposed measurements and policies are a starting point for monitoring the pace of AI development from outside the labs.
## (1) Measuring AI-led AI R&D
**Why measure AI-led R&D?** Frontier AI labs increasingly use AI to build future AI models. This process allows labs in democratic countries to develop more capable models more quickly and conduct more safety and testing on models before they are released to secure AI’s benefits while staying on the frontier. However, models accelerating their own development could make it more challenging for humans to understand or control these systems. It is therefore important to share these metrics to understand how close the world is to reaching [recursive self improvement](https://www.anthropic.com/institute/recursive-self-improvement) (a model fully autonomously building its successor).
**What we measured.** We built a prototype index of how much of Anthropic’s AI research and development (R&D) is performed by Claude, called the Anthropic R&D Automation Index. It’s built by cataloguing every kind of AI R&D work done at the company, rating how automated each task currently is, and aggregating those ratings.
**What we found.** To measure the extent to which AI is doing AI R&D at Anthropic, we use an [automation rating scale](https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d) developed by Epoch AI that measures “Automation Level,” or AL. It runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). In AL3, AI “collaborates”: it can do large chunks of work under close human direction. In AL4, AI “leads”: it can complete most of the task end-to-end from a high-level prompt, while the human supervises[1](https://www.anthropic.com/institute/measuring-pace-of-ai-development#footnote-1).
As of August 2026,
- Claude is not operating fully autonomously for any measured subset of AI R&D work.
- Claude “leads” 26% of Anthropic’s AI R&D work.
- The share of work at or above “AI collaborates” is above 90%.
Chart showing Claude now leads 26% of Anthropic's model R&D tasks, up from under 1% in February 2026.
**What any AI developer could report today.** Any frontier developer could publish these measures regularly, using a public methodology. This would enable the numbers to be compared over time, and potentially across labs.
Two obstacles stand in the way of cross-lab comparison on this type of reporting. First is the lack of a common methodology. Second, we’re using our own models to evaluate our systems, which could mean that the “judge” model could make the same kinds of errors as the model it is checking. To address these challenges, a developer’s measurements could be verified by a third party, or by other developers’ models, with guardrails in place to limit the exchange of competitively sensitive data. These measures could also become the trigger for stronger requirements, like a fixed testing window before a new model is used for further AI R&D.
## (2) Measuring oversight of AI agents
**Why measure oversight of agents?** Like other frontier developers, Anthropic employees increasingly delegate tasks to agents that work semi-autonomously for long stretches, and that delegate work to one another. As work becomes increasingly automated, from “AI collaborates” toward “AI leads,” agents could make more consequential decisions, such as which research direction to pursue next.
These decisions need safeguards to block harmful behavior from an individual agent, as well as an oversight regime to detect problems that might only arise across many agents. In our monitoring data to date, individual agents rarely misbehave. But when there are millions or billions of agents operating in the economy, even rare events can happen regularly.
**What we measured.** We’ve built a system that lets us oversee and intervene in actions AI agents take on Anthropic’s systems. Here, we consider three different metrics: *coverage,* which describes the share of an agent’s actions that pass through a monitor before or after execution; *review* *latency*, which is the time between an action and its review, first by an automated monitor and then by a human; and *escalation rate,* whichis the share of agent activities that are either blocked/redirected (in the case of online monitors) or flagged for further review (in the case of offline monitors).
**What we found.** As of August 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time in our most-used internal platform. These measurements cover this platform only. The actions of these agents are constrained by two kinds of monitors, summarized below:
| | | |
| --- | --- | --- |
| | **Online monitors** | **Offline monitors** |
| **Description** | Online monitors check the actions of agents *in real-time* and either block dangerous actions or steer agent misbehavior automatically. They’re built to address harms that would happen quickly and could not be undone, such as an agent copying model weights out of our systems. | Offline monitors read agent activities *after the fact*. They are built to help address dangerous behaviors that develop slowly and can be reversed, like signs of model misalignment. We use flagged activities to improve our mitigations, like classifiers and deterministic controls. If the situation caused a large imminent danger, we would shut down the dangerous parts of the agent’s deployment. |
| **Coverage** | 100% of these agents’ actions pass through the online monitor before they are executed, usually within seconds. | 100% of these agents’ actions are ingested by the monitor after they are taken. |
| **Review latency** | Automated review happens before an agent takes action; humans review any blocked actions within one week. | The monitor flags roughly 100,000 transcripts per week that are then reviewed by multiple classifiers for false-positives and initial oversight review. Most of the flags are resolved by this process, and the highest priority flags (approximately ~50 per week) are escalated to human review. |
| **Blocking / escalation rate** | We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor. | The monitor flags roughly one to two transcripts in every thousand for further review. |
**What an AI developer could report today.** Any developer running agents on its own research and engineering workloads could publish the same measures: *coverage* (the share of agent activity that is monitored), *review latency* (how soon flagged activities are reviewed), and *escalation rate* (what share of agent activities the monitors block or flag). Together, these measures would allow society to see whether oversight is keeping pace with the growing role of AI in AI R&D. We published all of these measurements in our recent [risk report](https://www.anthropic.com/aug-2026-risk-report).
Developers should share how often known agent misbehavior is caught by monitors, and should have third parties do this kind of testing independently. In the past, METR has [independently red-teamed our offline monitoring platform](https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf). As described above, we are now [setting up](https://darioamodei.com/post/we-must-pace-the-frontier) external third party evaluators at Anthropic.
## (3) Measuring compute allocation
**Why measure compute allocation?** Broadly speaking, AI developers use compute for building more powerful models, serving customers, and safety-focused work like [auditing a model’s “thoughts”](https://www.anthropic.com/research/natural-language-autoencoders), [training model organisms to study misalignment](https://www.anthropic.com/research/emergent-misalignment-reward-hacking), and [evaluating whether a model can be safely deployed](https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf). Understanding how AI developers allocate their compute can tell you where a developer is focusing its resources and how that focus changes over time.
Additionally, compute is among the most verifiable inputs to the AI R&D process, meaning that it could be a critical lever in a future pacing effort. A coordinated pacing effort could encourage companies to increase the compute allocated to safety across the industry and devote more resources to alignment, interpretability, safety testing, and evaluation.
**What we measured.** We examined a snapshot of how Anthropic used all of its compute from July 13 to July 20[2](https://www.anthropic.com/institute/measuring-pace-of-ai-development#footnote-2). To do that, we sorted every workload into a small number of categories, then asked how much of the compute going to AI R&D was safety work.
Safety research tends to use less compute than frontie | vertigoruntime | 5 | 0 |
| 🟠 reddit | Measurements for understanding the pace of AI development inside frontier labs singularity Retrieved article excerptOpen article · Retrieved 2026-09-17T21:22:07.370235+00:00 Measurements for
understanding the pace of AI development inside frontier labs
AI systems are becoming exponentially more powerful and have begun to [automate more](https://www.anthropic.com/institute/recursive-self-improvement) of the process of building themselves. As the world considers [slowing the pace of frontier AI development](https://darioamodei.com/post/we-must-pace-the-frontier), the public needs more information.
In this post, we lay out measurement tools that can illuminate three critical aspects of AI development:
1. [The extent to which AI is building the next version of itself, as opposed to being built by humans](https://www.anthropic.com/institute/measuring-pace-of-ai-development#1-measuring-ai-led-ai-rd)
2. [Our ability to oversee and intervene in actions that AI agents take on Anthropic’s systems](https://www.anthropic.com/institute/measuring-pace-of-ai-development#2-measuring-oversight-of-ai-agents)
3. [The resources that power the development of more capable models](https://www.anthropic.com/institute/measuring-pace-of-ai-development#3-measuring-compute-allocation)
We also provide a snapshot of these metrics from inside Anthropic. It’s important to note that we would expect these numbers to shift if there were coordination on pacing the frontier, as called for by Anthropic CEO Dario Amodei. We plan to embed independent third-party evaluators from multiple organizations at Anthropic, and give them access to internal processes, systems, and data comparable to what internal risk assessment teams have. These third parties will verify safety practices, report incidents, and monitor key metrics such as the ones in this piece.
We are reporting these measurements because they give the public, third parties, and governments better visibility into the pace of AI development inside frontier labs. For each measurement, we describe what we measured, what the measurement showed, and what it would take to publish these measurements regularly in a form others can verify. We share methodological details in the Appendix.
## Reasons to track these measurements
The measurements in this piece are focused on *how models are built.* By better understanding the production process of models, we have a better chance of correlating model inputs, like compute, with model outputs, like capabilities. They complement capability evaluations, which measure *what models can do*. We publish those separately through our [Responsible Scaling Policy](https://www.anthropic.com/responsible-scaling-policy) (RSP) risk reports, which include evidence on how much our models are accelerating AI R&D. In our policy proposal on advanced AI, the [Advanced AI Framework (AAIF)](https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf), we propose rules of the road for how any lab releases safe models, including transparency obligations that governments could require, such as risk reports. Together, these proposed measurements and policies are a starting point for monitoring the pace of AI development from outside the labs.
## (1) Measuring AI-led AI R&D
**Why measure AI-led R&D?** Frontier AI labs increasingly use AI to build future AI models. This process allows labs in democratic countries to develop more capable models more quickly and conduct more safety and testing on models before they are released to secure AI’s benefits while staying on the frontier. However, models accelerating their own development could make it more challenging for humans to understand or control these systems. It is therefore important to share these metrics to understand how close the world is to reaching [recursive self improvement](https://www.anthropic.com/institute/recursive-self-improvement) (a model fully autonomously building its successor).
**What we measured.** We built a prototype index of how much of Anthropic’s AI research and development (R&D) is performed by Claude, called the Anthropic R&D Automation Index. It’s built by cataloguing every kind of AI R&D work done at the company, rating how automated each task currently is, and aggregating those ratings.
**What we found.** To measure the extent to which AI is doing AI R&D at Anthropic, we use an [automation rating scale](https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d) developed by Epoch AI that measures “Automation Level,” or AL. It runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). In AL3, AI “collaborates”: it can do large chunks of work under close human direction. In AL4, AI “leads”: it can complete most of the task end-to-end from a high-level prompt, while the human supervises[1](https://www.anthropic.com/institute/measuring-pace-of-ai-development#footnote-1).
As of August 2026,
- Claude is not operating fully autonomously for any measured subset of AI R&D work.
- Claude “leads” 26% of Anthropic’s AI R&D work.
- The share of work at or above “AI collaborates” is above 90%.
Chart showing Claude now leads 26% of Anthropic's model R&D tasks, up from under 1% in February 2026.
**What any AI developer could report today.** Any frontier developer could publish these measures regularly, using a public methodology. This would enable the numbers to be compared over time, and potentially across labs.
Two obstacles stand in the way of cross-lab comparison on this type of reporting. First is the lack of a common methodology. Second, we’re using our own models to evaluate our systems, which could mean that the “judge” model could make the same kinds of errors as the model it is checking. To address these challenges, a developer’s measurements could be verified by a third party, or by other developers’ models, with guardrails in place to limit the exchange of competitively sensitive data. These measures could also become the trigger for stronger requirements, like a fixed testing window before a new model is used for further AI R&D.
## (2) Measuring oversight of AI agents
**Why measure oversight of agents?** Like other frontier developers, Anthropic employees increasingly delegate tasks to agents that work semi-autonomously for long stretches, and that delegate work to one another. As work becomes increasingly automated, from “AI collaborates” toward “AI leads,” agents could make more consequential decisions, such as which research direction to pursue next.
These decisions need safeguards to block harmful behavior from an individual agent, as well as an oversight regime to detect problems that might only arise across many agents. In our monitoring data to date, individual agents rarely misbehave. But when there are millions or billions of agents operating in the economy, even rare events can happen regularly.
**What we measured.** We’ve built a system that lets us oversee and intervene in actions AI agents take on Anthropic’s systems. Here, we consider three different metrics: *coverage,* which describes the share of an agent’s actions that pass through a monitor before or after execution; *review* *latency*, which is the time between an action and its review, first by an automated monitor and then by a human; and *escalation rate,* whichis the share of agent activities that are either blocked/redirected (in the case of online monitors) or flagged for further review (in the case of offline monitors).
**What we found.** As of August 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time in our most-used internal platform. These measurements cover this platform only. The actions of these agents are constrained by two kinds of monitors, summarized below:
| | | |
| --- | --- | --- |
| | **Online monitors** | **Offline monitors** |
| **Description** | Online monitors check the actions of agents *in real-time* and either block dangerous actions or steer agent misbehavior automatically. They’re built to address harms that would happen quickly and could not be undone, such as an agent copying model weights out of our systems. | Offline monitors read agent activities *after the fact*. They are built to help address dangerous behaviors that develop slowly and can be reversed, like signs of model misalignment. We use flagged activities to improve our mitigations, like classifiers and deterministic controls. If the situation caused a large imminent danger, we would shut down the dangerous parts of the agent’s deployment. |
| **Coverage** | 100% of these agents’ actions pass through the online monitor before they are executed, usually within seconds. | 100% of these agents’ actions are ingested by the monitor after they are taken. |
| **Review latency** | Automated review happens before an agent takes action; humans review any blocked actions within one week. | The monitor flags roughly 100,000 transcripts per week that are then reviewed by multiple classifiers for false-positives and initial oversight review. Most of the flags are resolved by this process, and the highest priority flags (approximately ~50 per week) are escalated to human review. |
| **Blocking / escalation rate** | We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor. | The monitor flags roughly one to two transcripts in every thousand for further review. |
**What an AI developer could report today.** Any developer running agents on its own research and engineering workloads could publish the same measures: *coverage* (the share of agent activity that is monitored), *review latency* (how soon flagged activities are reviewed), and *escalation rate* (what share of agent activities the monitors block or flag). Together, these measures would allow society to see whether oversight is keeping pace with the growing role of AI in AI R&D. We published all of these measurements in our recent [risk report](https://www.anthropic.com/aug-2026-risk-report).
Developers should share how often known agent misbehavior is caught by monitors, and should have third parties do this kind of testing independently. In the past, METR has [independently red-teamed our offline monitoring platform](https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf). As described above, we are now [setting up](https://darioamodei.com/post/we-must-pace-the-frontier) external third party evaluators at Anthropic.
## (3) Measuring compute allocation
**Why measure compute allocation?** Broadly speaking, AI developers use compute for building more powerful models, serving customers, and safety-focused work like [auditing a model’s “thoughts”](https://www.anthropic.com/research/natural-language-autoencoders), [training model organisms to study misalignment](https://www.anthropic.com/research/emergent-misalignment-reward-hacking), and [evaluating whether a model can be safely deployed](https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf). Understanding how AI developers allocate their compute can tell you where a developer is focusing its resources and how that focus changes over time.
Additionally, compute is among the most verifiable inputs to the AI R&D process, meaning that it could be a critical lever in a future pacing effort. A coordinated pacing effort could encourage companies to increase the compute allocated to safety across the industry and devote more resources to alignment, interpretability, safety testing, and evaluation.
**What we measured.** We examined a snapshot of how Anthropic used all of its compute from July 13 to July 20[2](https://www.anthropic.com/institute/measuring-pace-of-ai-development#footnote-2). To do that, we sorted every workload into a small number of categories, then asked how much of the compute going to AI R&D was safety work.
Safety research tends to use less compute than frontie | ResultBackground2450 | 10 | 1 |
| 🟧 echo.blog ⭐ | The Reddit title paraphrases Anthropic's oversight-of-AI-agents section: "We've built a system that lets us oversee and intervene in actions | Anthropic (co-authored by Marina Favaro and Phillie Wright, with editorial support from Santi Ruiz, Adam Farina, and Sarah Pollack; published by Anthropic PBC) | — | — |
| 🟠 reddit | Anthropic says its chatbot Claude is taking over the work of building its own successor ClaudeAI | Evening_Bicycle3113 | 427 | 93 |
| 🟧 hn | Anthropic says Claude now leads a quarter of work building its next AI models | dr_scully | 4 | 1 |
| 🟠 reddit | AI is now building AI: Anthropic says Claude leads 26% of its R&D work | Artificial Intelligence News ClaudeAI | 002Chris | 71 | 22 |
| 🟧 hn | Anthropic says its model Claude is helping to build the next version of itself | geox | 6 | 0 |
| 🟠 reddit | Anthropic says its model Claude is helping to build the next version of itself singularity | MrsSynchronie | 41 | 10 |
| 🟠 reddit | Claude itself is now leading 26% of the work building the next version of Claude. 7 months ago, it was 0%. "There are now 30,000 agents doing research and engineering work at Anthropic at any one time." ClaudeAI | AxomaticallyExtinct | 445 | 56 |
| 🟠 reddit | Claude now leads 26% of AI development at Anthropic singularity | vyxex | 166 | 29 |
| 🟧 hn | Measurements for understanding the pace of AI development inside frontier labs | gmays | 2 | 0 |
| 🟠 reddit | Anthropic's monitor blocks about 1 in 47,000 actions from its internal agents. That still works out to roughly 21,000 a month ClaudeAI | InterviewAsleep639 | 1 | 4 |
| 🟧 hn | Current ways of developing AI agents contribute to degradation of oversight | utiiiD | 4 | 1 |
| 🟧 hn | Automating eval design and hillclimbing with Claude | helloplanets | 2 | 0 |
| 🟧 hn | Automating eval design and hillclimbing with Claude | gmays | 3 | 0 |
| 🟧 hn | AI Agents Push Humans Out of the Loop | ibobev | 3 | 0 |
2026-09-29T23:19:42Z
Both new attachments are single-digit HN resubmissions of the two already-tracked side threads (Anthropic's eval-design/hillclimbing piece; the academic oversight-degradation work) — duplicates that leave the evidentiary structure unchanged: still one self-reported, unverified measurement, and the magnitude-valve flag remains an artifact of the Sept 18–19 spike with a static periphery (no new outlets, communities, or implementations). With engagement at <1 pt/h and zero comments 13 days past peak, the publication episode closes as absorbed — the claim banked as a dated receipt under the named re-open triggers, with any verification event (next index snapshot, third-party evaluators, Epoch/METR critique, rival-lab metrics) opening a fresh episode.
2026-09-29T20:53:11Z
evidence attached: hn.story.49897337 — shared external link with case evidence
2026-09-29T20:53:10Z
evidence attached: hn.story.49898332 — shared external link with case evidence
2026-09-29T07:52:51Z
The new Anthropic-side report of Claude automating eval design and hillclimbing is a concrete mechanistic exhibit of the delegated-work pattern — the lab showing how agent-executed R&D subtasks happen rather than only asserting the index — but it is the same source, so it neither breaks the single-actor evidentiary structure nor independently corroborates the index; the case still waits on third-party evaluators or the next snapshot, with engagement flatlined (~0.3 pts/h, zero comments/h) and the 74th peer percentile an artifact of a uniformly slow cohort.
2026-09-29T07:26:16Z
evidence attached: hn.story.49889097 — Anthropic-side report of Claude automating eval design and hillclimbing is concrete agent-executed R&D practice supporting the automation-index case's shift claim.
2026-09-27T16:32:04Z
The newly attached academic claim that current agent-development practices erode oversight is an independent counter-thread to Anthropic's 100%-monitor-coverage framing, but it does not address, validate, or contest the R&D Automation Index itself — the case remains one self-reported measurement awaiting third-party evaluators or the next snapshot. Engagement is fully decayed (~0.5 pts/h vs ~386 peak, 11 days old); the magnitude-valve flag reflects the Sept 18–19 cross-platform spike, not continuing expansion — recent additions are single-digit restatements — so heat stays low despite the valve and the hot sibling topics.
2026-09-27T16:26:17Z
evidence attached: hn.story.49867596 — Independent academic claim that current agent-development practices erode human oversight — direct context for the supervision side of scaling agent-executed R&D.
2026-09-24T14:33:17Z
grounded: converges/medium — Anthropic operationalizes Human Over the Loop at frontier-lab scale — Claude leads AL4 tasks end-to-end under supervision, nothing measured at AL5, monitors esc
2026-09-24T14:26:45Z
The episode is past its peak: current velocity is ~0.2 pts/h against a ~382 pt/h peak and 33rd peer percentile, and additions since the spike are single-digit-score restatements plus a derived arithmetic observation on blocking rates — the cross-platform spread that earned high heat has stopped expanding, so attention pricing drops to low without any change in belief. Evidentiary maturity is unchanged: still one self-reported Anthropic measurement with all coverage tracing to it, awaiting the promised third-party evaluators or the next index snapshot.
2026-09-24T13:50:39Z
evidence attached: reddit.post.1wp1dw4 — The same post restates the August 26%-of-R&D figure, adding circulation corroboration to the existing index case.
2026-09-21T18:22:31Z
The newly attached HN submission repeats the same primary publication and adds neither independent validation nor a research-throughput result, despite its substantive-evidence trigger. The supplied cross-platform magnitude signal still warrants high attention, but repeated coverage does not advance evidentiary maturity.
2026-09-21T18:21:58Z
evidence attached: hn.story.49790987 — shared external link with case evidence
2026-09-19T21:40:03Z
Scott’s up-vote confirms interest in the supervised-delegation and measurement implications, but the changed comments add no substantive evidence. The continuing cross-platform spread signal supports retaining high attention without treating repeated coverage as independent corroboration.
2026-09-19T06:23:24Z
The cross-platform spread signal and renewed engagement warrant attention even though the coverage still traces to one Anthropic report. This raises the story’s attention priority, not confidence in its measurement or claims of recursive research acceleration.
2026-09-18T19:57:33Z
The latest Reddit attachment links the same Anthropic publication and adds no independent measurement or implementation result. Repeated coverage does not strengthen the case beyond reported growth in supervised research delegation; neither autonomous successor development nor research-throughput gains are established.
2026-09-18T19:22:26Z
evidence attached: reddit.post.1wjypn0 — First-party-linked reporting provides additional visibility into Anthropic's claim that Claude performs 26% of its supervised AI R&D work.
2026-09-18T17:56:02Z
The latest Reddit attachment repeats Anthropic’s existing delegation and agent-count claims; its comments add neither independent validation nor implementation results. The headline’s “0%” also overstates the primary publication’s “under 1%” baseline, so this remains amplification of supervised delegation rather than new evidence of autonomous research or acceleration.
2026-09-18T17:23:48Z
evidence attached: reddit.post.1wjus8c — The high-engagement post directly relays Anthropic's reported rise to 26% agent-led R&D work, materially bearing on the open automation trend.
2026-09-18T16:42:50Z
The newly attached AP excerpt, quoted on Reddit, explicitly attributes the claim to Anthropic and adds no independent validation or implementation result. This remains evidence of reported growth in supervised research delegation, not demonstrated autonomous successor development or measured research acceleration.
2026-09-18T16:22:51Z
evidence attached: reddit.post.1wjufvp — shared external link with case evidence
2026-09-18T09:26:30Z
The newly attached headlines extend coverage of Anthropic’s self-report but supply no independent measurements; the attachment rationale describing AP coverage as corroboration exceeds the evidence provided. The defensible interpretation remains increased supervised delegation, not demonstrated autonomous successor development or measured research acceleration.
2026-09-18T09:21:27Z
evidence attached: hn.story.49751716 — Independent AP coverage corroborates the developing episode that Claude is materially assisting Anthropic’s next-model development.
2026-09-18T09:21:26Z
evidence attached: reddit.post.1wjknas — This provides outside coverage of Anthropic’s claim that Claude now performs a substantial share of its R&D work.
2026-09-18T01:38:55Z
The new HN item repeats Anthropic’s claim without supplying independent verification, despite the attachment rationale calling it corroboration. Discussion about faster research iteration highlights an existing unanswered question rather than demonstrating a productivity gain.
2026-09-18T01:21:38Z
evidence attached: hn.story.49748648 — This independently reported coverage directly corroborates Anthropic's claim that Claude now performs roughly a quarter of its AI R&D work.
2026-09-18T00:27:22Z
The successor-building coverage adds amplification, not independent validation or evidence of autonomous research. The supplied primary text defines supervised task leadership explicitly, making delegated execution—not recursive self-improvement or demonstrated research acceleration—the defensible interpretation.
2026-09-18T00:22:50Z
evidence attached: reddit.post.1wjaev2 — This coverage materially contextualizes Anthropic's claim that Claude is performing a growing share of successor-model R&D, but it is not independent corroboration.
2026-09-17T21:28:37Z
grounded: converges/medium — Anthropic’s reported expansion of Claude-led, human-supervised R&D converges with Scott’s Human Over the Loop position and extends the Claude-writing-itself arg
2026-09-17T21:22:35Z
case created — The three observations trace to one substantive measurement publication, distinct from Amodei's existing frontier-pacing commitment case and not independent corroboration of the reported results.