On 6 Sept 2026 OpenAI published 'Research acceleration: The view inside OpenAI', a self-audit declaring it had met the 'automated research intern' goal Sam Altman set in an October 2025 livestream: agents that carry out well-defined, days-long research tasks under human direction. By its mid-August measurements the research org runs 3.1 agent-workdays of agent effort per human workday (agent runtime only crossed above total human labor after June 2026), with median researchers spending $600+/day on inference and heavy use of 4-plus concurrent agent workflows, and it now targets a full 'automated AI researcher' by March 2028; Chief Scientist Jakub Pachocki writes he has 'a strong expectation' this progress sustains into recursive self-improvement, while the post itself concedes OpenAI 'does not yet know how to safely get all the way to aligned, full RSI.' The milestones are self-defined and self-measured: the 3.1 figure is aggregate agent runtime rather than researcher-equivalent output, high-level planning remains a minimal fraction of agent tokens, and independent validation of net research gains has so far come only as benchmark extrapolation (e.g., Vals' mid-2027 frontier-researcher trend, ahead of OpenAI's own target), not as output-side corroboration.
| source | object | author | score | comments |
| 🟠 reddit | OpenAI: AI agents now perform 3.1 researcher-workdays for every human researcher-workday, says it has reached “automated research intern” level, and expects “automated AI researcher” by March 2028 singularity | Neurogence | 449 | 77 |
| 🟧 echo.blog ⭐ | The Reddit excerpt quotes OpenAI reporting “3.1 agent-workdays of effort for every workday of human labor” as of mid-August, saying it has r | OpenAI | — | — |
| 🟠 reddit | From the Chief Scientist at OpenAI : An Alien Mind OpenAI | Bloated_Plaid | 390 | 74 |
| 🟠 reddit | OpenAI Chief Scientist: “Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement” singularity | Neurogence | 746 | 174 |
| 🟠 reddit | OpenAI say they already have an automated AI research intern and expect a full AI researcher by March 2028 which could lead to RSI. Is this legit or just IPO hype before the bubble pops? https://openai.com/index/research-acceleration-view-inside-openai/ singularity | ImmuneHack | 96 | 65 |
| 🟠 reddit | Research acceleration: The view inside OpenAI OpenAI | Psychological_Job614 | 118 | 15 |
| 🟠 reddit | Research acceleration: The view inside OpenAI singularity | Psychological_Job614 | 10 | 1 |
| 🟠 reddit | Areas where OpenAI researchers are spending AI tokens on. Probably indicative of how things will shape up in your office workspaces going ahead. singularity | No_Hovercraft6239 | 65 | 7 |
| 🟧 hn | Research acceleration: The view inside OpenAI | kyisaiah47 | 2 | 1 |
| 🟠 reddit | OpenAI putting safety first ... singularity | LatentSpaceLeaper | 281 | 43 |
| 🟠 reddit | 3 things that stood out in OpenAI’s Navier–Stokes post: training, live model upgrades, and 10,000 agents OpenAI | DataLearnerAI | 0 | 25 |
| 🟠 reddit | How GPT‑5.6 Sol helps run quantum computing experiments singularity | donutloop | 46 | 1 |
| 🟧 hn | How GPT‑5.6 Sol helps run quantum computing experiments | theanonymousone | 149 | 108 |
| 🟠 reddit | Astra Investigates SFT singularity | Leather_Area_2301 | 0 | 0 |
| 🟧 hn | OpenAI's apparent maths breakthrough raises profound questions | ijidak | 2 | 0 |
| 🟧 openai | Research acceleration: The view inside OpenAIRetrieved article excerptOpen article · Retrieved 2026-09-13T00:21:22.588025+00:00 September 6, 2026
[Research](https://openai.com/news/research/)[Publication](https://openai.com/research/index/publication/)[Safety](https://openai.com/news/safety-alignment/)
# Research acceleration: The view inside OpenAI
Loading…
Share
For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.
Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs.
We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements. According to our measurements, we have now reached the goal, [announced(opens in a new window)](https://x.com/sama/status/1983584366547829073?lang=en) last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.
Over the course of this year, OpenAI researchers’ daily work has changed substantially. Researchers are using coding agents throughout the day (often in concurrent sessions) and total usage is rapidly increasing, outpacing growth among other OpenAI teams. Researchers are contributing code faster and running more experiments. The ways researchers use agents are changing, too: agents are handling increasingly complex tasks, and succeeding at them more often. AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won’t keep pace with these specific metrics. But on the whole, these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress. People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.
If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. It can bring down the cost of advanced intelligence so that people worldwide can benefit. We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.
These are reasons to develop useful automated research capabilities, but they do not mean that rapid RSI is necessarily an outcome we should pursue. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks.
We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.
After the recent Hugging Face incident, we [put this commitment into action](https://openai.com/index/pacing-model-development-cyber-capabilities/), pausing reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded coverage of our monitoring systems. This did not halt all research: some workloads resumed under stronger controls, while others remained paused. We have raised our safety and alignment standards and moved safety work deeper into the model lifecycle, requiring stronger evidence of aligned behavior throughout all of training.
Today we are providing a detailed snapshot of how agentic systems have contributed to our progress toward RSI in recent months. Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary. By sharing these early results and the methods behind them, we aim to inform the public, encourage a norm of public disclosure, and help the field move toward shared standards of measurement.
Ultimately, as we wrote in our [frontier policy blueprint](https://openai.com/index/frontier-safety-blueprint/), we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve, while balancing the need to protect security and proprietary information.
## 1. Coding agents are reshaping daily work for OpenAI researchers
At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. The 90th percentile user in our research organization now uses more than $7,000 of tokens per day.
View methods
View methods
Before June 2026, total agent runtime across the research organization was still below that of total human labor. That has since changed. In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.
View methods
Another way of looking at this is to understand how many researchers use highly concurrent workflows (e.g., running 4 or more agents simultaneously). As shown below, this number is increasing. These figures include the daily peaks of both agents started directly by the user and subagents created downstream from those the user launched directly.
View methods
## 2. Researchers are writing more code and running more experiments
Much of AI research can be seen as a labor-intensive process with the goal of integrating a new improvement to model intelligence or performance into one of our core models. The process depends on many steps, and capabilities advance when all the steps go right together: Researchers have to design new improvements, write evaluations to judge model performance, write infrastructure to test these improvements at scale, catch bugs as well as unsafe or misaligned behavior during training, and integrate winning ideas into a core training run. A failure at any part of the research process can constrain the entire loop.
Writing code and running experiments are two major activities that researchers do as part of their work, and we see evidence that these processes are accelerating.
View methods
These data points are relatively easy to measure, but can be hard to interpret. As automation progresses, the tasks which are *least* automatable will take on a larger share of researcher effort and will become the important bottlenecks to future progress. Compute is another gating factor for progress, and may become more important over time as other bottlenecks diminish.
Through 2026, the number of experiments per active experimenter has increased, with August 2026 being an all-time high since tracking began in Jan 2025. This is correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.
View methods
## 3. The work researchers use agents for is changing
Both qualitative impressions and internal data indicate that the mix of tasks researchers delegate to coding agents is changing, with delegation of higher level and longer-horizon tasks becoming more common over time.
To get a clearer picture of this trend, we analyzed recent usage in the research organization using a [recently published taxonomy(opens in a new window)](https://epoch.ai/gradient-updates/toward-an-onet-for-ai-rnd) of the different kinds of work that are part of the AI R&D lifecycle, developed by Epoch AI. This taxonomy, inspired by the longstanding O\*NET system for classifying all kinds of work, is specifically tailored to frontier AI R&D, and breaks the process down into six main phases:
1. Decide: what to work on, what to continue, where to allocate
2. Design: research ideas and engineering specs
3. Build: code and datasets
4. Run: training/eval runs, hardware, serving
5. Analyze: experiments, models, deployment, external work
6. Communicate: findings, feedback, status, decisions
Below, we classify coding agent tokens under this taxonomy.
View methods
View methods
We see that all categories of research activities have increased between January and August 2026. In January, the dominant category was research and infrastructure code. This category has expanded, but we also see notable increases in additional categories, especially technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.
Anecdotally, colleagues report that coding agents excel at troubleshooting internal research infrastructure, which addresses one meaningful bottleneck to research progress. Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely, to focus on making other system improvements instead.
Here, we plot the number of top-level posts per day to one of the main internal channels where researchers seek technical support from other teams. To our knowledge, the channel’s decrease in activity has not been offset by queries shifting to another technical support channel run by humans. The decline in traffic aligns with this broader shift.
View methods
We can also study whether coding agents are succeeding at the tasks researchers request. Using an agentic classifier, we find that from January to July, success rates generally increased across several difficulty buckets (proxied as the estimated time a human would take to complete the task) on tasks we can find a ground truth outcome for. However, agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions.
Success rates on researcher tasks have increased over time. Graph excludes classifications where the outcome was uncertain and points with <50 sessions or <50 unique users.
View methods
Task success and intervention rate from Jan to July, broken out by time horizon. Excludes classifications where the outcome was uncertain.
View methods
## 4. Pacing model development
Progress toward more capable systems for safe and beneficial AGI will also depend on the safeguards needed for such work. Our assessment of the needed safeguards may change as we learn more about the risks.
[As we have described,](https://openai.com/index/pacing-model-development-cyber-capabilities/) we have recently updated our standards for monitoring, alignment, and security. Here, we show how recent restrictions have affected one aspect of research activity.
\*The majority of | OpenAI | — | — |
| 🟧 openai | An Alien MindRetrieved article excerptOpen article · Retrieved 2026-09-12T00:21:20.823282+00:00 September 6, 2026
[Safety](https://openai.com/news/safety-alignment/)[Research](https://openai.com/news/research/)
# An Alien Mind
By: Jakub Pachocki, Chief Scientist at OpenAI
Loading…
Share
In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
## Intellect we don’t fully understand
At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects[1](https://openai.com/index/an-alien-mind/#citation-bottom-1). As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.
There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.
And, in line with [Ray Kurzweil’s predictions from the end of the XXth century(opens in a new window)](https://www.thekurzweillibrary.com/the-coming-merging-of-mind-and-machine), we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.
AI is *grown* more than *designed* - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.
The study of deep learning-based AI is largely an experimental science. We [put a lot of effort](https://openai.com/index/gpt-4-research/#predictable-scaling) into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are *experiments*, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.
This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
## Teaching machines to love
Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of *alignment* - getting the AI to “try to do the right thing” by human standards.
For the purpose of organizing practical research directions, I find it useful to distinguish *goal alignment* and *value alignment*.
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an [instruction hierarchy](https://openai.com/index/the-instruction-hierarchy/), or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the [intent(opens in a new window)](https://ai-alignment.com/clarifying-ai-alignment-cec47cd69dd6) and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.
The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
There are two major classes of currently practically employed methods for alignment training.
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model’s actions are evaluated (usually by AI) for being consistent with a given preference model, “spec” or “constitution”, and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model’s ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.
The second approach seeks to leverage the model’s ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an ‘aligned’ part of the pretraining distribution, as in, for example, the [persona selection model(opens in a new window)](https://www.anthropic.com/research/persona-selection-model). The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.
We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
## Monitoring generalization
We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.
OpenAI’s primary bet here has been [chain-of-thought monitoring(opens in a new window)](https://arxiv.org/abs/2507.11473). It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.
We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to [hide the chain of thought](https://openai.com/index/learning-to-reason-with-llms/#hiding-the-chains-of-thought), to protect it from supervision pressure in the long term[2](https://openai.com/index/an-alien-mind/#citation-bottom-2). In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishi | OpenAI | — | — |
| 🟧 openai | How GPT-5.6 Sol helps run quantum computing experimentsRetrieved article excerptOpen article · Retrieved 2026-09-12T00:21:23.553902+00:00 September 8, 2026
Applied AI
# How GPT‑5.6 Sol helps run quantum computing experiments
Connecting GPT‑5.6 Sol to laboratory software to run and refine routine measurements on quantum chips freed Beatriz Yankelevich to focus on experiment design and data analysis.
[Read the technical case study(opens in a new window)](https://cdn.openai.com/pdf/case-study-agentic-calibration-of-superconducting-qubits.pdf)
Loading…
Share
Quantum computing is an emerging technology that uses the unique properties of quantum mechanics to process information. It could one day better simulate complex materials and molecules. Unlike conventional processors, quantum processors are built with quantum bits, or qubits. Preparing and running qubit experiments can take months and require hundreds to thousands of preliminary measurements—work that AI is poised to help with.
Beatriz Yankelevich, a graduate student in MIT’s Engineering Quantum Systems Group (EQuS), used GPT‑5.6 Sol, harnessed to Codex, to explore whether AI could streamline her experimental workflow. The MIT group studies superconducting qubits, which are cooled to near absolute zero inside specialized devices called dilution refrigerators. These qubits perform operations quickly, are precisely controlled using microwave signals, and can be made using familiar manufacturing techniques and arranged on a chip.
Once a superconducting qubit chip has been fabricated, packaged, and cooled, researchers interact with it entirely through software, making Yankelevich’s experiments a natural testbed for AI agents. Connecting Codex to the lab software that coordinates experiments allowed it to run measurements, analyze the results, and decide what to try next. Yankelevich found that GPT‑5.6 Sol could often complete routine measurement workflows autonomously, saving her significant amounts of time and allowing experiments to run without constant supervision. This freed her to spend more time on analyzing results, designing experiments, and planning out the next steps in her research.
A packaged qubit chip beside an open dilution refrigerator with cabling that connects to the chip.
*A packaged qubit chip (left) sits inside an open dilution refrigerator (right). CREDIT: EQuS group*
## Coordinating interdependent measurements
Superconducting qubits are often called artificial atoms because, like atoms, they can only occupy specific energy levels. Microwave pulses move qubits between these levels and probe their quantum state. Researchers design and calibrate the pulse sequences sent to the chip, then digitize and analyse the returning signals. These measurements reveal each qubit’s resonance frequencies, which allows researchers to accurately control the qubit; how long the qubit retains quantum information; and the settings needed to perform computations.
Calibrating qubits requires a series of interdependent measurements, with each result shaping what happens next. Qubit properties can occasionally drift, and unexpected physical behavior can cause inconsistent results. Experienced researchers can recognize these changes and adapt when they occur. This combination of software control, repeated measurements, and adaptive decision-making also makes qubit calibration a compelling use case for AI agents.
Yankelevich tested GPT‑5.6 Sol’s ability to run measurements on an uncalibrated six-qubit chip, one of a standard type that EQuS routinely uses to benchmark its fabrication process. She provided Codex with measurement-specific skills explaining how to run and evaluate each experiment. Using these skills and the chip’s design targets, GPT‑5.6 Sol chose measurement parameters, operated the hardware, analyzed the resulting data, and then either refined the measurement or saved the result for use in the next measurement.
When the signals were clear, Codex completed a standard sequence of measurements with little researcher intervention. It identified the qubit’s transition frequencies, calibrated the pulses used to control and read it, and determined how long the qubit retained quantum information.
Two q1 calibration plots show a fitted resonance dip by frequency and a fitted Rabi oscillation by drive power.
*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*
Two q1 calibration plots show fitted relaxation and coherence measurements over pulse duration.
*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*
Four q1 readout-calibration plots show overlap, IQ clusters, and ground- and excited-state histograms.
*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*
GPT‑5.6 Sol had more difficulty when experimental signals were weak or noisy. In those cases, it took longer to find suitable measurement parameters and sometimes needed guidance from an experienced researcher. The results suggest that current agents can handle clearly defined experimental workflows, but interpreting ambiguous physical results remains a challenge.
EQuS fabricates many of these standard chips, each of which can take a researcher several days to characterize. The group now regularly uses agents to handle routine measurements, freeing researchers to focus on other work.
“I can have agents running measurements for many hours overnight or while I’m working in the cleanroom,” Yankelevich said. “I can check in from my phone, see what they’ve done, and steer them if something needs fixing or if I want to explore a different direction.”
Two GPT-5.6 Sol frequency-amplitude calibration screenshots show Rabi and Ramsey calibration results and plots against a pink gradient.
*An excerpted GPT‑5.6 Sol chain-of-thought from a calibration run. CREDIT: EQuS group*
## Working alongside researchers
The immediate advantage is that Codex agents can help researchers make steady progress on experimental analysis and measurements without constant supervision. Experienced researchers may still be able to identify the best calibration settings faster than current AI models. But by saving time previously spent on monitoring every step of the calibration process, researchers can focus on other work.
Routine chip characterization follows a relatively well-defined workflow. For novel experiments, Yankelevich assigns Codex agents narrower experimental goals while drawing more heavily on their ability to write, modify, and test new code for control, analysis, and simulation. Connecting agents directly to the lab lets the group revise code, test it against real measurements, and complete longer stretches of work autonomously.
“I’ve built infrastructure to guide agents through several parts of my work—measurement, theory, and chip design—and now it’s really starting to pay off,” Yankelevich said. “I can have multiple agents working on different problems at once, and I spend most of my time on higher-level work—interpreting results, devising experiments, planning next steps for the agents, reading, and writing.”
- 2026
- Codex
## Author
OpenAI
## Keep reading
[View all](https://openai.com/news/)
How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules — card image
[How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
Applied AISep 10, 2026](https://openai.com/index/using-codex-chatgpt-to-search-for-new-antimicrobials/)
The builder’s guide to GPT-5.6 — Card image — Neutral Option 075
[The builder’s guide to GPT‑5.6
Applied AIAug 13, 2026](https://openai.com/index/builders-guide-to-gpt-5-6/)
Derya Unutmaz card image
[How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
Applied AIJun 23, 2026](https://openai.com/index/gpt-5-immunology-mystery/) | OpenAI | — | — |
| 🟠 reddit | How GPT‑5.6 Sol helps run quantum computing experiments OpenAI | truecakesnake | 1 | 0 |
| 🟧 hn | Can AI agents conduct open-ended AI research? | Betelbuddy | 3 | 0 |
| 🟠 reddit | RSI is not happening [R] MachineLearning | we_are_mammals | 285 | 160 |
| 🟧 hn | Inside OpenAI’s agentic software factory | gfortaine | 1 | 0 |
| 🟧 hn | OpenAI's Agentic Software Factory | ishener | 3 | 0 |
| 🟠 reddit | From The Information -- incredible (if true) reports of the AI models assisting with AI training at OpenAI singularity | BrennusSokol | 224 | 37 |
| 🟠 reddit | OpenAI says the AIs themselves are now doing most of the work training the next AIs OpenAI | Puzzleheaded-King584 | 65 | 11 |
| 🟧 hn | People Training OpenAI's AI Fired for Using AI to Train the AI | pier25 | 80 | 57 |
| 🟠 reddit | OpenAI researchers might be spending >$4m/day on tokens (at API prices) OpenAI | abtin | 334 | 48 |
| 🟧 hn | OpenAI researchers might be spending >$4M/day on tokens (at API prices) | abtinf | 2 | 1 |
| 🟧 hn | Recursive self-improvement of AI research agents | handfuloflight | 2 | 1 |
| 🟠 reddit | Two different methods, same date: Vals’ measured autonomous-R&D trend now points to frontier-level AI researchers by July 2027—the exact month AI 2027 projected singularity | 141_1337 | 61 | 8 |
| 🟠 reddit | Can AI automate AI R&D yet? singularity | Proper_Actuary2907 | 33 | 19 |
| 🟠 reddit | OpenAI Researcher: It’s Really Not as Easy to Train Models to Do AI R&D as It Is to Do Math singularity | Neurogence | 122 | 37 |
| 🟧 hn | Can AI automate AI R&D yet? | merksittich | 78 | 61 |
2026-10-10T03:20:38Z
No material developments since last reprice: the Epoch AI article (now on HN) and an unverified OpenAI researcher tweet merely reinforce the existing two-skeptical-vs-one-bullish independent-evidence split. Engagement is at floor (~0.8 pts/h) despite a 64th-percentile peer reading that reflects the spent September peak, not current expansion.
2026-10-10T01:44:11Z
evidence attached: hn.story.50027257 — shared external link with case evidence
2026-10-08T00:30:49Z
The new evidence is a 5-point tweet attributed to an OpenAI researcher conceding that automating AI R&D is much harder than math — sentiment that merely restates what the NeurIPS-replication study and Epoch AI already established, with unverified attribution and no new finding. The case's meaning is unchanged: first-party runtime claims vs. a structured two-skeptical-vs-one-bullish split of independent checks, held in quiet monitoring until a live trigger (Epoch article-level confirmation, The Information article text, an OpenAI output-side snapshot) fires.
2026-10-08T00:26:04Z
evidence attached: reddit.post.1x0cell — An OpenAI researcher's insider caveat that automating AI R&D is much harder than math directly contextualizes the automated-researcher claims in that case.
2026-10-07T22:38:29Z
Epoch AI's 'Can AI automate AI R&D yet?' gives the case a second independent output-side evaluation, converging with the NeurIPS-replication study's negative result and now squaring off against the Vals timeline extrapolation: the independent evidence is structured (two direct output evaluations skeptical vs one benchmark trend bullish) rather than single-line, sharpening the live question from 'is there any independent check?' to 'do runtime/activity metrics or direct output evaluations predict the March 2028 trajectory?'. Material on substance, not attention — the magnitude-valve cross-platform spread reflects the spent September peak; current flow is ~2.5 pts/h at the 40th peer percentile.
2026-10-07T22:34:02Z
evidence attached: reddit.post.1x05ds9 — Independent Epoch AI publication evaluating whether AI can automate AI R&D is exactly the external yardstick needed to re-judge OpenAI's automated-researcher trajectory.
2026-10-02T12:52:36Z
The velocity spike and comment drift all trace to the Vals trend post already priced in at the last reprice — 18.7x a near-zero peer baseline but only ~4.7 pts/h absolute, decayed to zero within a day — repetitive amplification with no new facts, sources or implementations. The case's meaning is unchanged (corroborated on OpenAI's first-party claims plus Vals' independent timeline extrapolation, still lacking output-side validation); with engagement at floor and the September cross-platform spread fully spent, attention cools to low while the re-entry triggers stay live.
2026-10-01T05:01:24Z
grounded: converges/high — OpenAI's own audit concedes the 3.1 figure is aggregate runtime rather than researcher-equivalent output — the largest live instantiation of the exact seam Scot
2026-10-01T04:54:15Z
The case gains its first genuine independent second line: Vals' benchmark-trend measurement independently projects frontier-level AI researchers by mid-2027 — converging with AI 2027 and running ahead of OpenAI's own March 2028 target — moving the timeline question from OpenAI-asserted to independently measured-forecast. State upgrades to corroborated on substance, but with engagement near floor this is an evidence upgrade, not an attention event.
2026-10-01T04:25:25Z
evidence attached: reddit.post.1wunwqw — Independent Vals trend measurement converging on mid-2027 frontier AI researchers is material external evidence on the automated-researcher timeline the case tracks.
2026-09-25T22:01:54Z
grounded: converges/high — Converges on the exact seam Scott's canon already holds: OpenAI's own audit concedes the 3.1 figure is aggregate runtime rather than output (over half of succes
2026-09-25T21:54:31Z
The five velocity alerts all trace to a late second engagement wave on the already-known $4M/day token-spend post (score ~218→281, brief ~12 pts/h against a floor baseline) plus minor comment drift on two other old items — repetitive cost-discussion, not new facts, sources or implementations. The case's meaning is unchanged: first-party claims stand, independent validation of net research gains is still absent, The Information's 'AI training AI' line remains headline-only, so the post-peak cooling from the last reprice holds.
2026-09-23T23:22:09Z
The episode has moved from breaking claim to established-claims-plus-contested-substance: OpenAI's first-party disclosures are stable, and this week's additions (token-spend estimate, contractor firings, academic RSI counter-evidence, The Information lead) are contextual color rather than new facts about the trajectory. Engagement has cooled hard from its peak (288→5.5 pts/h) and remaining chatter is derivative, so attention temperature drops even though the underlying question — independently validated net research gains — stays open.
2026-09-23T21:46:37Z
evidence attached: hn.story.49819845 — Independent academic work on recursive self-improvement of research agents materially contextualises the automated-researcher trajectory in the open case.
2026-09-23T21:46:36Z
evidence attached: hn.story.49819381 — Estimated >$4M/day internal token spend is cost-side context for OpenAI's claim of agent-executed research running at scale.
2026-09-23T21:46:36Z
evidence attached: reddit.post.1woct12 — Third-party estimate of >$4m/day internal token spend quantifies the scale of agent usage behind the automated-researcher claims, though it is unverified.
2026-09-22T14:23:55Z
evidence attached: hn.story.49800953 — The reported contractor firings materially contextualize OpenAI’s claimed shift toward AI-executed R&D by showing possible limits on human use of AI in training work.
2026-09-22T12:23:32Z
evidence attached: reddit.post.1wn7o75 — This is redundant coverage of the claim that OpenAI is using AI systems to perform substantial work on future AI development.
2026-09-21T17:33:40Z
The new Reddit post points to possible reporting by The Information, extending the episode’s apparent media reach, but supplies neither the article nor concrete findings sufficient to establish independent corroboration. Renewed coverage alongside the broad cross-platform spread signal warrants high attention, not greater confidence in measured research acceleration.
2026-09-21T17:25:08Z
evidence attached: reddit.post.1wmhvqf — Independent reporting echoed by a high-engagement community post supports the open hypothesis that OpenAI is already using agents materially in frontier-model research.
2026-09-21T14:16:06Z
The latest “agentic software factory” link is another low-engagement, headline-only derivative and adds no implementation detail or measured outcome. The episode remains broadly distributed across platforms, but its periphery is no longer expanding fast enough to justify hours-level attention.
2026-09-21T08:21:38Z
evidence attached: hn.story.49784462 — shared external link with case evidence
2026-09-19T21:45:36Z
Scott’s explicit up-vote raises the case’s personal relevance, particularly the distinction between agent runtime and validated research gains, but adds no evidence for the capability claims. Broad cross-platform attention still warrants high heat; neither the vote nor marginal engagement growth changes maturity.
2026-09-19T07:22:44Z
The cross-platform spread signal and renewed attention to the open-ended-research critique raise the episode’s attention priority, not confidence in either research acceleration or its rebuttal. No new methods, implementation results or independent validation accompany this spike; supervised task automation remains distinct from demonstrated net research gains.
2026-09-17T19:22:12Z
The new “agentic software factory” attachment supplies only a headline, not implementation details or measured outcomes, so it adds no corroboration of research acceleration. The case remains about demonstrated company-reported supervised task automation versus unresolved net research gains and autonomous research capability.
2026-09-17T19:21:49Z
evidence attached: hn.story.49745280 — Reporting on OpenAI's agentic software factory could materially contextualize claims that internal agents are already executing substantial research and engineering work.
2026-09-14T18:44:53Z
The new Reddit account gives the open-ended-research paper a reported negative direction and a sketch of its evaluation, but the truncated secondary summary does not establish its results or applicability to OpenAI’s internal agents. It sharpens the distinction between supervised task execution and open-ended research without disproving the intern milestone or settling the 2028 forecast.
2026-09-14T18:22:33Z
evidence attached: reddit.post.1wgazy4 — The paper offers relevant independent evidence against current agents already performing open-ended ML research, directly constraining claims about automated research and recursive improvement.
2026-09-14T17:49:33Z
The new open-ended-research attachment supplies only a title, not the paper’s methods or results, so it cannot yet corroborate or contradict OpenAI’s capability claims. The distinction remains supervised research-task automation versus independently demonstrated net research acceleration.
2026-09-14T17:24:41Z
evidence attached: hn.story.49700395 — The paper directly bears on whether agents can perform open-ended AI research, a key external test of the capability trajectory in the open case.
2026-09-13T00:22:45Z
The revised primary account specifies that safety restrictions included pausing RL training for deployment-bound models, with some workloads subsequently resuming under stronger controls while others remained paused. This makes safeguards a concrete constraint on research throughput, but supplies no new validation of net research acceleration or the automated-researcher timeline.
2026-09-12T14:26:41Z
The new attachment only repeats the already assessed MIT quantum-lab case study; it adds neither an implementation result nor independent validation. The case remains evidence of supervised research-task automation, not demonstrated net frontier-R&D acceleration, and repetitive coverage does not warrant continued medium heat.
2026-09-12T14:22:23Z
evidence attached: reddit.post.1wedmxu — shared external link with case evidence
2026-09-12T00:28:36Z
The full primary source confirms that OpenAI explicitly claims the research-intern milestone, but its intervention data sharpen the distinction between supervised task execution and autonomous research productivity. A named MIT lab implementation adds concrete evidence for agents running bounded experimental workflows, not independent validation of OpenAI’s frontier-R&D gains or recursive self-improvement.
2026-09-12T00:22:13Z
evidence attached: openai.article.1f112fc7530a72cea21528f2 — shared external link with case evidence
2026-09-12T00:22:13Z
evidence attached: openai.article.1f0643794f0c78a697d9ef7c — shared external link with case evidence
2026-09-12T00:22:13Z
evidence attached: openai.article.be17d270d974aa3c372f03a6 — shared external link with case evidence
2026-09-10T05:22:35Z
The refreshed discussion identifies checkpoint provenance as a potential problem for live model upgrades, but offers no evidence that it occurred in OpenAI’s campaign. This adds an evaluation question, not a deployment change or validation of research gains; the previously reported automation claims remain unresolved.
2026-09-09T22:33:52Z
The new maths-breakthrough headline broadens coverage of the already reported swarm campaign, but the supplied evidence contains no expert assessment, proof validation, or operational detail. It does not independently corroborate OpenAI’s claimed research gains or materially change the previously alerted deployment testimony.
2026-09-09T22:22:43Z
evidence attached: hn.story.49635093 — The reported mathematics breakthrough could materially contextualize whether OpenAI's claimed agent-assisted research progress is producing substantive results.
2026-09-09T20:38:47Z
The refreshed quantum-experiment discussion adds reactions about promotion and economics, not evidence of agent responsibilities or experimental outcomes. OpenAI’s reported research automation remains a consequential strategy signal, but this delta neither validates productive research gains nor materially changes the previously alerted swarm campaign.
2026-09-09T17:27:18Z
The refreshed quantum-experiment discussion adds no evidence of agent contributions, experimental outcomes, or gains over conventional automation. OpenAI’s reported deployments remain meaningful research-strategy testimony, but this delta neither validates research productivity nor materially changes the previously alerted swarm campaign.
2026-09-09T16:32:55Z
Refreshed quantum-experiment comments add speculation about announcement timing and economics, not evidence of agent contributions or experimental outcomes. The reported research deployments remain meaningful company-attributed testimony, but this delta does not strengthen capability validation or materially change the previously alerted swarm campaign.
2026-09-09T14:34:20Z
The new manuscript link adds a third-party claim of AI-authored investigation, not corroboration of OpenAI’s internal research automation: the supplied excerpt contains no findings, execution evidence, or independent assessment. The deployment testimony remains meaningful, but this attachment does not close the gap between agent activity and demonstrated research gains.
2026-09-09T14:24:00Z
evidence attached: reddit.post.1wbmh6c — The released forensic manuscript is a concrete artifact bearing on whether frontier models can execute substantive research tasks rather than merely assist with them.
2026-09-09T13:35:53Z
The refreshed discussion adds reactions and analogies, not evidence of agent responsibilities, experimental outcomes, or gains over conventional automation. OpenAI’s reported research-agent deployments remain consequential testimony, but this delta neither independently validates the intern milestone nor changes the previously alerted swarm campaign.
2026-09-09T12:34:19Z
Refreshed swarm and quantum-experiment comments remain amplification and speculation, not new operational evidence or research validation. The reported deployments retain their significance, but this delta does not establish useful research gains or warrant revisiting the existing alert.
2026-09-09T11:30:34Z
The refreshed quantum-experiment discussion adds skepticism and analogies, not evidence of agent responsibilities, experimental outcomes, or gains over conventional automation. The reported research-agent deployments remain consequential testimony, but this delta does not independently validate capability or research productivity.
2026-09-09T10:26:03Z
The refreshed quantum-experiment discussion adds speculation about announcement timing and economics, not experimental results or evidence of what agents contributed beyond conventional automation. OpenAI’s reported research-agent deployments remain meaningful testimony, but this delta does not strengthen or rebut the claimed capability gains.
2026-09-09T09:28:46Z
The new quantum-experiment discussion raises a useful baseline question—what agents add beyond conventional laboratory automation—but offers only an unverified anecdote, not comparative evidence. Neither this discussion nor the refreshed swarm commentary changes the substantive deployment testimony or establishes measured research gains.
2026-09-09T08:31:16Z
The HN attachment repeats the already-linked quantum-experiment report without adding its contents, so it broadens circulation rather than corroborating capability. The case remains a meaningful OpenAI-attributed research-automation signal, with no new evidence connecting agent activity to useful research output.
2026-09-09T08:22:20Z
evidence attached: hn.story.49622561 — shared external link with case evidence
2026-09-09T06:23:44Z
The refreshed discussion supplies no operational evidence for live model upgrades or measured research gains, and the quantum-experiment attachment remains title-only. The substantive signal is still OpenAI-attributed research-agent deployment testimony; this delta neither strengthens nor rebuts the claimed capability gains.
2026-09-09T05:27:45Z
The quantum-experiment attachment points toward another research-agent application, but its title-only evidence does not establish agent responsibilities, experimental outcomes, or a new deployment capability; the attachment rationale overstates what is supplied. The broader research-automation claim remains meaningful company-attributed testimony, without new evidence connecting agent effort to useful research output.
2026-09-09T05:22:20Z
evidence attached: reddit.post.1wbc5hr — OpenAI's first-party Codex quantum-experiment report is concrete evidence relevant to whether its agents can execute meaningful research workflows.
2026-09-09T03:25:46Z
The refreshed comments question continuity and cost but provide no operational evidence about live model upgrades or research outcomes. The reported swarm campaign remains substantive company-attributed testimony; this delta neither strengthens nor rebuts the claimed research-automation gains.
2026-09-09T01:22:45Z
The new paraphrase introduces a potentially useful workflow detail—upgrading research agents while model training continues—but the supplied excerpt does not substantiate that mechanism or its benefits. This extends the reported deployment story without independently validating research productivity, intern-level capability, or recursive improvement.
2026-09-09T01:22:04Z
evidence attached: reddit.post.1wb6h7h — The post highlights OpenAI’s reported concurrent training, model upgrades, and large-scale agent use in its research workflow.
2026-09-09T00:26:12Z
Refreshed comments repeat the swarm claim and add an unsupported runtime figure, not a substantiated deployment change or research result. The campaign remains meaningful company-attributed testimony, but neither it nor aggregate agent runtime independently establishes research productivity or recursive improvement.
2026-09-08T23:24:22Z
The refreshed comments add an unverified cached-internet detail and repeat claims about additional swarms, without establishing a new deployment event or research outcome. The large-agent campaign remains substantive company-attributed testimony, but does not independently validate mathematical correctness, research productivity, or recursive improvement.
2026-09-08T21:46:03Z
Refreshed comments amplify the previously reported swarm campaign but add no substantiated deployment detail, cost measurement, or research validation. The campaign remains a substantive company-attributed research-automation claim, distinct from proof of mathematical correctness or productive recursive improvement.
2026-09-08T20:30:13Z
The new company-attributed quotation extends the strategy signal from aggregate agent runtime to a concrete claimed research campaign involving roughly 10,000 concurrent agents and a Navier–Stokes resolution. This is substantive deployment testimony, but not independent corroboration of mathematical correctness, researcher-equivalent productivity, or recursive improvement; the claimed safeguards likewise lack operational detail.
2026-09-08T20:23:18Z
evidence attached: reddit.post.1wayu26 — The cited OpenAI report adds context about large agent swarms being used in frontier mathematical research.
2026-09-08T13:36:02Z
The new HN attachment adds coverage but supplies no article contents or independent research results, so its attachment rationale overstates corroboration. OpenAI’s reported internal deployment remains consequential, but the quoted intern milestone and runtime ratio still do not establish productive research gains or recursive improvement.
2026-09-08T13:22:58Z
evidence attached: hn.story.49609686 — This independent write-up materially corroborates and contextualizes OpenAI's claims about agents accelerating internal research workflows.
2026-09-08T06:30:30Z
The refreshed discussion adds generic agent-use anecdotes and reactions, not a new OpenAI deployment disclosure or measured research outcome. The reported internal automation remains a meaningful strategy signal, but the intern milestone is still quoted testimony and aggregate agent runtime does not establish useful research gains.
2026-09-07T22:33:41Z
The refreshed comments add endorsements and criticism, not a new deployment disclosure or measured research outcome. OpenAI’s reported research-automation strategy remains meaningful, but repeated testimony does not independently validate intern-level capability or convert agent runtime into useful research output.
2026-09-07T19:39:30Z
The refreshed discussion is repetitive amplification of the existing disclosure, not a new research result or deployment change. OpenAI’s reported research-automation strategy remains meaningful, but the quoted intern milestone remains unvalidated and agent runtime does not establish productive research gains.
2026-09-07T18:23:36Z
The refreshed chief-scientist discussions add endorsements and safety reactions, not new internal results or a deployment change. OpenAI’s reported research-automation strategy remains meaningful, but repeated company-attributed testimony does not independently validate intern-level capability or translate agent runtime into useful research output.
2026-09-07T16:28:54Z
The refreshed comments add generic agent-use anecdotes and reactions, not evidence of OpenAI research outcomes or a deployment change. The company-attributed automation strategy remains meaningful, but neither the echoed intern milestone nor aggregate agent runtime establishes researcher-equivalent productivity; review should now focus on substantive results rather than thread activity.
2026-09-07T14:36:12Z
The refreshed discussion adds endorsements and skepticism, not new evidence of research output or a deployment change. OpenAI’s reported internal automation remains a meaningful strategy signal, but repeated testimony neither independently validates the intern milestone nor turns agent runtime into demonstrated research productivity.
2026-09-07T12:34:38Z
The refreshed chief-scientist discussion adds endorsements and reactions, not new internal results or deployment details. OpenAI’s reported research-automation strategy remains consequential, but the echoed intern milestone and agent-runtime metric still do not demonstrate productive research gains or achieved recursive improvement.
2026-09-07T09:29:26Z
The refreshed discussion remains amplification of OpenAI’s reported research-automation strategy, with no new deployment detail or research result. The intern milestone remains quoted testimony, and aggregate agent runtime does not establish productive research gains; further hourly review of these same threads is unlikely to add value.
2026-09-07T08:26:53Z
Refreshed comments repeat reactions to the existing disclosure rather than add deployment details or research results. OpenAI’s reported internal research automation remains a consequential strategy signal, but the echoed intern milestone and agent-runtime ratio still do not establish useful research gains or achieved recursive improvement.
2026-09-07T07:33:36Z
The newly attached post links the same OpenAI disclosure without supplying token-allocation details or a separate research result; it is amplification, not additional corroboration. The meaningful signal remains OpenAI’s reported internal research-agent deployment and strategy, with productive research gains and intern-level capability still unvalidated.
2026-09-07T07:22:36Z
evidence attached: reddit.post.1w9l6o9 — The linked OpenAI research post provides additional first-party context for the open case on agents performing substantial automated research work.
2026-09-07T04:25:50Z
The refreshed discussion is repetitive amplification, not a new research result or deployment change. OpenAI’s reported internal automation remains a consequential strategy signal, but the quoted intern milestone lacks independent validation and the runtime ratio does not establish productive research gains.
2026-09-07T03:23:03Z
The refreshed discussion adds reactions to OpenAI’s safety stance and research ambitions, not new deployment or research-output evidence. The substantive signal remains company-attributed research automation and strategy; repeated quotations do not validate intern-level capability or convert agent runtime into productive research gains.
2026-09-07T02:33:35Z
The refreshed discussion sharpens the distinction between agent activity and useful research output but supplies no new evidence resolving it. OpenAI’s reported internal deployment remains a consequential strategy signal, not demonstrated researcher-equivalent productivity or recursive improvement.
2026-09-07T01:29:41Z
The refreshed discussion adds reactions rather than a new research result or deployment detail. OpenAI’s reported internal automation remains a meaningful strategy signal, but repeated quotations do not independently validate intern-level capability, productive research gains, or recursive improvement.
2026-09-07T00:22:50Z
Refreshed discussion adds neither a deployment change nor demonstrated research gains; the meaningful signal remains OpenAI’s reported internal research-automation strategy. The intern milestone is quoted testimony, while agent-workdays measure effort rather than useful output or achieved recursive improvement.
2026-09-06T23:51:35Z
Only comment/engagement churn on already-known evidence; no independent validation of research-output or milestone achievement beyond the original OpenAI report. Remains a single-source claim of runtime metrics and roadmap aspiration.
2026-09-06T22:33:16Z
Only comment/engagement refresh on already-known evidence; no new capability validation, independent corroboration, or research-output evidence has emerged since last look. Case remains a single-source (OpenAI first-party) report of runtime metrics and roadmap aspiration.
2026-09-06T22:09:13Z
Comment activity on existing evidence without new capability validation or independent corroboration. Case remains a single-source report of internal deployment metrics and roadmap aspirations. No progression.
2026-09-06T20:43:54Z
New evidence are duplicate links to the same primary source; no independent validation or new capability evidence. Case remains a single-source report of internal deployment metrics and roadmap aspirations. No progression towards corroboration.
2026-09-06T20:27:15Z
evidence attached: reddit.post.1w96mpn — Duplicate of observation 158301, same primary source for the open case.
2026-09-06T20:27:15Z
evidence attached: reddit.post.1w96m2w — Primary source from OpenAI on internal agent research acceleration, directly supports the open case.
2026-09-06T19:29:19Z
The chief-scientist excerpts connect the reported internal agent deployment to an explicit expectation of recursive self-improvement, making this a research-strategy signal rather than merely a runtime headline. They remain company-attributed testimony, not independent capability validation: agent-workdays do not establish productive research output or achievement of the stated milestones.
2026-09-06T19:22:51Z
evidence attached: reddit.post.1w94ol9 — Direct discussion of OpenAI’s claimed automated research-intern capability and its path toward autonomous AI research.
2026-09-06T19:22:51Z
evidence attached: reddit.post.1w959f0 — This independently surfaced quotation reinforces OpenAI's explicit claim that internal agent progress may sustain recursive self-improvement.
2026-09-06T19:22:51Z
evidence attached: reddit.post.1w940xw — A first-party chief-scientist essay provides direct context for OpenAI's stated push toward AI-driven research and recursive self-improvement.
2026-09-06T18:31:02Z
The refreshed discussion adds speculation, not deployment details or research-output evidence, so this remains a single-line report rather than corroborated capability. The reported milestone still merits tracking, but the 3.1 agent-workdays ratio establishes neither researcher-equivalent productivity nor autonomous research success.
2026-09-06T17:26:50Z
This remains a consequential but single-line report of internal research-agent deployment: the echo reconstructs the same Reddit quotation, not independent corroboration. The new discussion adds no evidence of achieved intern capability or productive research output; the 3.1 figure remains an effort measure.
2026-09-06T17:26:02Z
grounded: converges/medium — OpenAI’s reported deployment of research agents under human supervision converges with Scott’s Persistent Intelligence Layer position—machine-scale attention wi
2026-09-06T17:23:25Z
case created — A linked first-party account presents a distinct internal research-automation milestone and dated roadmap, although its effort ratio does not establish equivalent productive human research output.