2026-10-11 16:37 UTC

OpenAI claims its internal agents have reached “automated research intern” capability, with 3.1 agent-workdays of effort per human workday, and are progressing toward an automated AI researcher by March 2028, potentially shifting frontier-model R&D toward agent-executed research.

state: corroboratedheat: lowuncertainty: highconvergesscott: highresearch-agents recursive-improvement coding-agentsOpenAI
Surfaced 2026-09-21T17:33:40Z — The Reddit excerpt quotes OpenAI reporting “3.1 agent-workdays of effort for every workday of human labor” as of mid-August, saying it has r — The new Reddit post points to possible reporting by The Information, extending the episode’s apparent media reach, but supplies neither the article nor concrete findings sufficient to establish independent corroboration. Renewed coverage alongside the broad cross-platform spread signal warrants high attention, not greater confidence in measured research acceleration.

What is this?

On 6 Sept 2026 OpenAI published 'Research acceleration: The view inside OpenAI', a self-audit declaring it had met the 'automated research intern' goal Sam Altman set in an October 2025 livestream: agents that carry out well-defined, days-long research tasks under human direction. By its mid-August measurements the research org runs 3.1 agent-workdays of agent effort per human workday (agent runtime only crossed above total human labor after June 2026), with median researchers spending $600+/day on inference and heavy use of 4-plus concurrent agent workflows, and it now targets a full 'automated AI researcher' by March 2028; Chief Scientist Jakub Pachocki writes he has 'a strong expectation' this progress sustains into recursive self-improvement, while the post itself concedes OpenAI 'does not yet know how to safely get all the way to aligned, full RSI.' The milestones are self-defined and self-measured: the 3.1 figure is aggregate agent runtime rather than researcher-equivalent output, high-level planning remains a minimal fraction of agent tokens, and independent validation of net research gains has so far come only as benchmark extrapolation (e.g., Vals' mid-2027 frontier-researcher trend, ahead of OpenAI's own target), not as output-side corroboration.

Why it matters to Scott

OpenAI's own audit concedes the 3.1 figure is aggregate runtime rather than researcher-equivalent output — the largest live instantiation of the exact seam Scott's Mature Token Law and Capability-Denominated Return already hold (tokens are fuel, not the score), and a dated-receipts publishing opportunity: the primary source itself admits the vanity-metric problem (>$4M/day spend, >50% of successful multi-hour tasks needing human intervention, compute confounds). The new Vals trend adds the first genuinely independent second line on the timeline, but it is still benchmark extrapolation — measuring activity, not the output-side conversion audit his frameworks demand — while the reported intervention rates bear directly on his long-running-agents and barbell-supervision doctrine.
ip:framework.the-mature-token-lawip:concept.capability-denominated-returnip:concept.ai-unit-economicsip:concept.self-improving-loopsip:framework.long-running-agentsip:concept.barbell-supervisionradar:concept.ai-research-agentsradar:concept.automated-researchradar:concept.recursive-self-improvementradar:concept.self-improving-agentsradar:anthropic-rnd-automation-indexradar:concept.inference-economics
queries asked of Scott's wikis
  • multi-agent orchestration concurrent coding agent harness
  • agent runtime versus output productivity measurement
  • recursive self-improvement automated researcher timeline positions
  • long-running agent task intervention and supervision rates
  • agent inference cost economics token spend scaling
  • research artifact verification provenance agent-maintained wiki

Measured heat

now 0 pts/hpeak 45 pts/hcomments 0/hpeers p50momentum: steady4 platformsage 848h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-06 17:23 (minted)⭐ origin echo-reconstructedThe Reddit excerpt quotes OpenAI reporting “3.1 agent-workdays of effort for every workday of human labor” as of mid-August, saying it has r
OpenAI on blog (echo) · attributed from reddit.post.1w914qw · published time unknown
—
09-06 08:00first on openai · published · lag ?Research acceleration: The view inside OpenAI
OpenAI
—
09-06 16:42first on r/singularity · published · lag ?OpenAI: AI agents now perform 3.1 researcher-workdays for every human researcher-workday, says it has reached “automated research intern” level, and expects “automated AI researcher” by March 2028
Neurogence
—
09-06 18:28first on r/OpenAI · published · lag ?From the Chief Scientist at OpenAI : An Alien Mind
Bloated_Plaid
—
09-08 12:52first on hacker news · published · lag ?Research acceleration: The view inside OpenAI
kyisaiah47
—
09-14 18:03first on r/MachineLearning · published · lag ?RSI is not happening [R]
we_are_mammals
—
09-06 16:42amplified on r/singularityreddit.post.1w914qw
Neurogence
peak 449 · 77 comments · 10% of case engagement
09-06 18:28amplified on r/OpenAIreddit.post.1w940xw
Bloated_Plaid
peak 390 · 74 comments · 9% of case engagement
09-06 18:52amplified on r/singularityreddit.post.1w94ol9
ImmuneHack
peak 96 · 65 comments · 3% of case engagement
09-06 19:13amplified on r/singularity 👑reddit.post.1w959f0
Neurogence
peak 746 · 174 comments · 18% of case engagement
09-06 20:03amplified on r/OpenAIreddit.post.1w96m2w
Psychological_Job614
peak 119 · 15 comments · 3% of case engagement
09-06 20:03amplified on r/singularityreddit.post.1w96mpn
Psychological_Job614
peak 10 · 1 comments · 0% of case engagement
22 more amplifiers in ainews.case_chain
09-06 17:20our radar first saw it · lag ?discovery anchor: reddit.post.1w914qw—
09-19 07:22reached heat=high · lag ? · via ledger——
pace: p98 vs 519 stories at the 720h mark (now 848h old) — ahead of moonshot-claude-routing-allegation (1.0x), behind gpt-6-astra-safety-controls (1.0x)

Evidence (33) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditOpenAI: AI agents now perform 3.1 researcher-workdays for every human researcher-workday, says it has reached “automated research intern” level, and expects “automated AI researcher” by March 2028
singularity
Neurogence44977
🟧 echo.blog ⭐The Reddit excerpt quotes OpenAI reporting “3.1 agent-workdays of effort for every workday of human labor” as of mid-August, saying it has rOpenAI——
🟠 redditFrom the Chief Scientist at OpenAI : An Alien Mind
OpenAI
Bloated_Plaid39074
🟠 redditOpenAI Chief Scientist: “Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement”
singularity
Neurogence746174
🟠 redditOpenAI say they already have an automated AI research intern and expect a full AI researcher by March 2028 which could lead to RSI. Is this legit or just IPO hype before the bubble pops? https://openai.com/index/research-acceleration-view-inside-openai/
singularity
ImmuneHack9665
🟠 redditResearch acceleration: The view inside OpenAI
OpenAI
Psychological_Job61411815
🟠 redditResearch acceleration: The view inside OpenAI
singularity
Psychological_Job614101
🟠 redditAreas where OpenAI researchers are spending AI tokens on. Probably indicative of how things will shape up in your office workspaces going ahead.
singularity
No_Hovercraft6239657
🟧 hnResearch acceleration: The view inside OpenAIkyisaiah4721
🟠 redditOpenAI putting safety first ...
singularity
LatentSpaceLeaper28143
🟠 reddit3 things that stood out in OpenAI’s Navier–Stokes post: training, live model upgrades, and 10,000 agents
OpenAI
DataLearnerAI025
🟠 redditHow GPT‑5.6 Sol helps run quantum computing experiments
singularity
donutloop461
🟧 hnHow GPT‑5.6 Sol helps run quantum computing experimentstheanonymousone149108
🟠 redditAstra Investigates SFT
singularity
Leather_Area_230100
🟧 hnOpenAI's apparent maths breakthrough raises profound questionsijidak20
🟧 openaiResearch acceleration: The view inside OpenAI
Retrieved article excerpt

Open article · Retrieved 2026-09-13T00:21:22.588025+00:00

September 6, 2026

[Research](https://openai.com/news/research/)[Publication](https://openai.com/research/index/publication/)[Safety](https://openai.com/news/safety-alignment/)

# Research acceleration: The view inside OpenAI

Loading…

Share

For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.

Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs.

We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements. According to our measurements, we have now reached the goal, [announced⁠(opens in a new window)](https://x.com/sama/status/1983584366547829073?lang=en) last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.

Over the course of this year, OpenAI researchers’ daily work has changed substantially. Researchers are using coding agents throughout the day (often in concurrent sessions) and total usage is rapidly increasing, outpacing growth among other OpenAI teams. Researchers are contributing code faster and running more experiments. The ways researchers use agents are changing, too: agents are handling increasingly complex tasks, and succeeding at them more often. AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won’t keep pace with these specific metrics. But on the whole, these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress. People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.

If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. It can bring down the cost of advanced intelligence so that people worldwide can benefit. We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.

These are reasons to develop useful automated research capabilities, but they do not mean that rapid RSI is necessarily an outcome we should pursue. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks.

We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.

After the recent Hugging Face incident, we [put this commitment into action⁠](https://openai.com/index/pacing-model-development-cyber-capabilities/), pausing reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded coverage of our monitoring systems. This did not halt all research: some workloads resumed under stronger controls, while others remained paused. We have raised our safety and alignment standards and moved safety work deeper into the model lifecycle, requiring stronger evidence of aligned behavior throughout all of training.

Today we are providing a detailed snapshot of how agentic systems have contributed to our progress toward RSI in recent months. Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary. By sharing these early results and the methods behind them, we aim to inform the public, encourage a norm of public disclosure, and help the field move toward shared standards of measurement.

Ultimately, as we wrote in our [frontier policy blueprint⁠](https://openai.com/index/frontier-safety-blueprint/), we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve, while balancing the need to protect security and proprietary information.

## 1. Coding agents are reshaping daily work for OpenAI researchers

At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. The 90th percentile user in our research organization now uses more than $7,000 of tokens per day.

View methods

View methods

Before June 2026, total agent runtime across the research organization was still below that of total human labor. That has since changed. In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.

View methods

Another way of looking at this is to understand how many researchers use highly concurrent workflows (e.g., running 4 or more agents simultaneously). As shown below, this number is increasing. These figures include the daily peaks of both agents started directly by the user and subagents created downstream from those the user launched directly.

View methods

## 2. Researchers are writing more code and running more experiments

Much of AI research can be seen as a labor-intensive process with the goal of integrating a new improvement to model intelligence or performance into one of our core models. The process depends on many steps, and capabilities advance when all the steps go right together: Researchers have to design new improvements, write evaluations to judge model performance, write infrastructure to test these improvements at scale, catch bugs as well as unsafe or misaligned behavior during training, and integrate winning ideas into a core training run. A failure at any part of the research process can constrain the entire loop.

Writing code and running experiments are two major activities that researchers do as part of their work, and we see evidence that these processes are accelerating.

View methods

These data points are relatively easy to measure, but can be hard to interpret. As automation progresses, the tasks which are *least* automatable will take on a larger share of researcher effort and will become the important bottlenecks to future progress. Compute is another gating factor for progress, and may become more important over time as other bottlenecks diminish.

Through 2026, the number of experiments per active experimenter has increased, with August 2026 being an all-time high since tracking began in Jan 2025. This is correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.

View methods

## 3. The work researchers use agents for is changing

Both qualitative impressions and internal data indicate that the mix of tasks researchers delegate to coding agents is changing, with delegation of higher level and longer-horizon tasks becoming more common over time.

To get a clearer picture of this trend, we analyzed recent usage in the research organization using a [recently published taxonomy⁠(opens in a new window)](https://epoch.ai/gradient-updates/toward-an-onet-for-ai-rnd) of the different kinds of work that are part of the AI R&D lifecycle, developed by Epoch AI. This taxonomy, inspired by the longstanding O\*NET system for classifying all kinds of work, is specifically tailored to frontier AI R&D, and breaks the process down into six main phases:

1. Decide: what to work on, what to continue, where to allocate
2. Design: research ideas and engineering specs
3. Build: code and datasets
4. Run: training/eval runs, hardware, serving
5. Analyze: experiments, models, deployment, external work
6. Communicate: findings, feedback, status, decisions

Below, we classify coding agent tokens under this taxonomy.

View methods

View methods

We see that all categories of research activities have increased between January and August 2026. In January, the dominant category was research and infrastructure code. This category has expanded, but we also see notable increases in additional categories, especially technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.

Anecdotally, colleagues report that coding agents excel at troubleshooting internal research infrastructure, which addresses one meaningful bottleneck to research progress. Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely, to focus on making other system improvements instead.

Here, we plot the number of top-level posts per day to one of the main internal channels where researchers seek technical support from other teams. To our knowledge, the channel’s decrease in activity has not been offset by queries shifting to another technical support channel run by humans. The decline in traffic aligns with this broader shift.

View methods

We can also study whether coding agents are succeeding at the tasks researchers request. Using an agentic classifier, we find that from January to July, success rates generally increased across several difficulty buckets (proxied as the estimated time a human would take to complete the task) on tasks we can find a ground truth outcome for. However, agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions.

Success rates on researcher tasks have increased over time. Graph excludes classifications where the outcome was uncertain and points with <50 sessions or <50 unique users.

View methods

Task success and intervention rate from Jan to July, broken out by time horizon. Excludes classifications where the outcome was uncertain.

View methods

## 4. Pacing model development

Progress toward more capable systems for safe and beneficial AGI will also depend on the safeguards needed for such work. Our assessment of the needed safeguards may change as we learn more about the risks.

[As we have described,⁠](https://openai.com/index/pacing-model-development-cyber-capabilities/) we have recently updated our standards for monitoring, alignment, and security. Here, we show how recent restrictions have affected one aspect of research activity.

\*The majority of 
OpenAI——
🟧 openaiAn Alien Mind
Retrieved article excerpt

Open article · Retrieved 2026-09-12T00:21:20.823282+00:00

September 6, 2026

[Safety](https://openai.com/news/safety-alignment/)[Research](https://openai.com/news/research/)

# An Alien Mind

By: Jakub Pachocki, Chief Scientist at OpenAI

Loading…

Share

In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.

Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.

A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.

This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.

## Intellect we don’t fully understand

At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects[1](https://openai.com/index/an-alien-mind/#citation-bottom-1). As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.

There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.

And, in line with [Ray Kurzweil’s predictions from the end of the XXth century⁠(opens in a new window)](https://www.thekurzweillibrary.com/the-coming-merging-of-mind-and-machine), we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.

AI is *grown* more than *designed* - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.

The study of deep learning-based AI is largely an experimental science. We [put a lot of effort⁠](https://openai.com/index/gpt-4-research/#predictable-scaling) into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are *experiments*, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.

This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.

The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.

## Teaching machines to love

Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of *alignment* - getting the AI to “try to do the right thing” by human standards.

For the purpose of organizing practical research directions, I find it useful to distinguish *goal alignment* and *value alignment*.

Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an [instruction hierarchy⁠](https://openai.com/index/the-instruction-hierarchy/), or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.

Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.

Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the [intent⁠(opens in a new window)](https://ai-alignment.com/clarifying-ai-alignment-cec47cd69dd6) and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.

The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.

There are two major classes of currently practically employed methods for alignment training.

The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model’s actions are evaluated (usually by AI) for being consistent with a given preference model, “spec” or “constitution”, and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model’s ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.

The second approach seeks to leverage the model’s ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an ‘aligned’ part of the pretraining distribution, as in, for example, the [persona selection model⁠(opens in a new window)](https://www.anthropic.com/research/persona-selection-model). The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.

We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.

## Monitoring generalization

We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.

OpenAI’s primary bet here has been [chain-of-thought monitoring⁠(opens in a new window)](https://arxiv.org/abs/2507.11473). It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.

We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to [hide the chain of thought⁠](https://openai.com/index/learning-to-reason-with-llms/#hiding-the-chains-of-thought), to protect it from supervision pressure in the long term[2](https://openai.com/index/an-alien-mind/#citation-bottom-2). In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.

This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishi
OpenAI——
🟧 openaiHow GPT-5.6 Sol helps run quantum computing experiments
Retrieved article excerpt

Open article · Retrieved 2026-09-12T00:21:23.553902+00:00

September 8, 2026

Applied AI

# How GPT‑5.6 Sol helps run quantum computing experiments

Connecting GPT‑5.6 Sol to laboratory software to run and refine routine measurements on quantum chips freed Beatriz Yankelevich to focus on experiment design and data analysis.

[Read the technical case study(opens in a new window)](https://cdn.openai.com/pdf/case-study-agentic-calibration-of-superconducting-qubits.pdf)

Loading…

Share

Quantum computing is an emerging technology that uses the unique properties of quantum mechanics to process information. It could one day better simulate complex materials and molecules. Unlike conventional processors, quantum processors are built with quantum bits, or qubits. Preparing and running qubit experiments can take months and require hundreds to thousands of preliminary measurements—work that AI is poised to help with.

Beatriz Yankelevich, a graduate student in MIT’s Engineering Quantum Systems Group (EQuS), used GPT‑5.6 Sol, harnessed to Codex, to explore whether AI could streamline her experimental workflow. The MIT group studies superconducting qubits, which are cooled to near absolute zero inside specialized devices called dilution refrigerators. These qubits perform operations quickly, are precisely controlled using microwave signals, and can be made using familiar manufacturing techniques and arranged on a chip.

Once a superconducting qubit chip has been fabricated, packaged, and cooled, researchers interact with it entirely through software, making Yankelevich’s experiments a natural testbed for AI agents. Connecting Codex to the lab software that coordinates experiments allowed it to run measurements, analyze the results, and decide what to try next. Yankelevich found that GPT‑5.6 Sol could often complete routine measurement workflows autonomously, saving her significant amounts of time and allowing experiments to run without constant supervision. This freed her to spend more time on analyzing results, designing experiments, and planning out the next steps in her research.

A packaged qubit chip beside an open dilution refrigerator with cabling that connects to the chip.

*A packaged qubit chip (left) sits inside an open dilution refrigerator (right). CREDIT: EQuS group*

## Coordinating interdependent measurements

Superconducting qubits are often called artificial atoms because, like atoms, they can only occupy specific energy levels. Microwave pulses move qubits between these levels and probe their quantum state. Researchers design and calibrate the pulse sequences sent to the chip, then digitize and analyse the returning signals. These measurements reveal each qubit’s resonance frequencies, which allows researchers to accurately control the qubit; how long the qubit retains quantum information; and the settings needed to perform computations.

Calibrating qubits requires a series of interdependent measurements, with each result shaping what happens next. Qubit properties can occasionally drift, and unexpected physical behavior can cause inconsistent results. Experienced researchers can recognize these changes and adapt when they occur. This combination of software control, repeated measurements, and adaptive decision-making also makes qubit calibration a compelling use case for AI agents.

Yankelevich tested GPT‑5.6 Sol’s ability to run measurements on an uncalibrated six-qubit chip, one of a standard type that EQuS routinely uses to benchmark its fabrication process. She provided Codex with measurement-specific skills explaining how to run and evaluate each experiment. Using these skills and the chip’s design targets, GPT‑5.6 Sol chose measurement parameters, operated the hardware, analyzed the resulting data, and then either refined the measurement or saved the result for use in the next measurement.

When the signals were clear, Codex completed a standard sequence of measurements with little researcher intervention. It identified the qubit’s transition frequencies, calibrated the pulses used to control and read it, and determined how long the qubit retained quantum information.

Two q1 calibration plots show a fitted resonance dip by frequency and a fitted Rabi oscillation by drive power.

*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*

Two q1 calibration plots show fitted relaxation and coherence measurements over pulse duration.

*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*

Four q1 readout-calibration plots show overlap, IQ clusters, and ground- and excited-state histograms.

*A set of calibration measurements for one qubit, completed autonomously by GPT‑5.6 Sol. CREDIT: EQuS group*

GPT‑5.6 Sol had more difficulty when experimental signals were weak or noisy. In those cases, it took longer to find suitable measurement parameters and sometimes needed guidance from an experienced researcher. The results suggest that current agents can handle clearly defined experimental workflows, but interpreting ambiguous physical results remains a challenge.

EQuS fabricates many of these standard chips, each of which can take a researcher several days to characterize. The group now regularly uses agents to handle routine measurements, freeing researchers to focus on other work.

“I can have agents running measurements for many hours overnight or while I’m working in the cleanroom,” Yankelevich said. “I can check in from my phone, see what they’ve done, and steer them if something needs fixing or if I want to explore a different direction.”

Two GPT-5.6 Sol frequency-amplitude calibration screenshots show Rabi and Ramsey calibration results and plots against a pink gradient.

*An excerpted GPT‑5.6 Sol chain-of-thought from a calibration run. CREDIT: EQuS group*

## Working alongside researchers

The immediate advantage is that Codex agents can help researchers make steady progress on experimental analysis and measurements without constant supervision. Experienced researchers may still be able to identify the best calibration settings faster than current AI models. But by saving time previously spent on monitoring every step of the calibration process, researchers can focus on other work.

Routine chip characterization follows a relatively well-defined workflow. For novel experiments, Yankelevich assigns Codex agents narrower experimental goals while drawing more heavily on their ability to write, modify, and test new code for control, analysis, and simulation. Connecting agents directly to the lab lets the group revise code, test it against real measurements, and complete longer stretches of work autonomously.

“I’ve built infrastructure to guide agents through several parts of my work—measurement, theory, and chip design—and now it’s really starting to pay off,” Yankelevich said. “I can have multiple agents working on different problems at once, and I spend most of my time on higher-level work—interpreting results, devising experiments, planning next steps for the agents, reading, and writing.”

- 2026
- Codex

## Author

OpenAI

## Keep reading

[View all](https://openai.com/news/)

How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules — card image

[How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

Applied AISep 10, 2026](https://openai.com/index/using-codex-chatgpt-to-search-for-new-antimicrobials/)

The builder’s guide to GPT-5.6 — Card image — Neutral Option 075

[The builder’s guide to GPT‑5.6

Applied AIAug 13, 2026](https://openai.com/index/builders-guide-to-gpt-5-6/)

Derya Unutmaz card image

[How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery

Applied AIJun 23, 2026](https://openai.com/index/gpt-5-immunology-mystery/)
OpenAI——
🟠 redditHow GPT‑5.6 Sol helps run quantum computing experiments
OpenAI
truecakesnake10
🟧 hnCan AI agents conduct open-ended AI research?Betelbuddy30
🟠 redditRSI is not happening [R]
MachineLearning
we_are_mammals285160
🟧 hnInside OpenAI’s agentic software factorygfortaine10
🟧 hnOpenAI's Agentic Software Factoryishener30
🟠 redditFrom The Information -- incredible (if true) reports of the AI models assisting with AI training at OpenAI
singularity
BrennusSokol22437
🟠 redditOpenAI says the AIs themselves are now doing most of the work training the next AIs
OpenAI
Puzzleheaded-King5846511
🟧 hnPeople Training OpenAI's AI Fired for Using AI to Train the AIpier258057
🟠 redditOpenAI researchers might be spending >$4m/day on tokens (at API prices)
OpenAI
abtin33448
🟧 hnOpenAI researchers might be spending >$4M/day on tokens (at API prices)abtinf21
🟧 hnRecursive self-improvement of AI research agentshandfuloflight21
🟠 redditTwo different methods, same date: Vals’ measured autonomous-R&D trend now points to frontier-level AI researchers by July 2027—the exact month AI 2027 projected
singularity
141_1337618
🟠 redditCan AI automate AI R&D yet?
singularity
Proper_Actuary29073319
🟠 redditOpenAI Researcher: It’s Really Not as Easy to Train Models to Do AI R&D as It Is to Do Math
singularity
Neurogence12237
🟧 hnCan AI automate AI R&D yet?merksittich7861

Interpretation history

Decision trace