2026-10-11 17:15 UTC

OpenAI says uncertainty over whether Astra reaches its Critical cybersecurity threshold caused a two-week pause in deployment-bound reinforcement learning and additional monitoring compute costs, making frontier cyber risk a direct constraint on model development and deployment.

state: resolvedheat: lowuncertainty: mediumknownscott: highfrontier-cyber-safety training-monitoring inference-economicsOpenAI

What is this?

OpenAI says preliminary evaluations could not rule out that its upcoming Astra model reaches the “Critical” cybersecurity capability level in its Preparedness Framework. In response, it paused deployment-oriented reinforcement-learning training for two weeks while hardening and red-teaming research environments, and imposed stricter controls including isolation, restricted tool/network access, sandboxing, weight protection, and universal monitoring for agentic use. A secondary report estimates that monitoring adds roughly 20% to monitored inference compute, but that figure is not established in the supplied first-party snippets.

Why it matters to Scott

The radar already tracks this development in “OpenAI will translate its cyber-capability pacing framework into concrete model evaluations, development gates, or release controls” and “OpenAI will implement substantive frontier-model cyber-safety or deployment controls.” The reported RL pause and containment measures directly bear on Scott’s Gate Criteria, Earned Autonomy, and SiloOS work by showing a frontier lab allowing capability evidence to block advancement and trigger structural isolation, although the claimed monitoring-compute premium remains unconfirmed.
ip:framework.gate-criteria-frameworkip:concept.earned-autonomyip:concept.runtime-containmentdev:project.silo-osradar:openai-cyber-capability-pacingradar:openai-frontier-cyber-controlsradar:openai-training-slowdown
queries asked of Scott's wikis
  • frontier capabilities as a constraint on model scaling
  • runtime monitoring economics for coding agents
  • agent sandboxing and restricted tool access
  • capability thresholds tied to deployment gates
  • security controls for model training environments
  • compute tax of continuous agent monitoring

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (12) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditOpenAI paused RL training for two weeks and added a monitoring compute cost over Astra's cyber tier reading
OpenAI
coursiv_106
🟧 echo.blog ⭐The Reddit post synthesizes OpenAI’s Aug. 18 first-party post: “a two-week pause in reinforcement learning (RL) training” for deployment modOpenAI——
🟠 redditOpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities
OpenAI
wiredmagazine11232
🟠 redditPath to Astra: critical capabilities and frontier safeguards
singularity
Ok_Display_315921170
🟧 hnOpenAI: Path to Astra: critical capabilities and frontier safeguardsjithinraj17094
🟧 hnOpenAI to Restrict Astra Model After Rating It 'Critical' Cyber Riskbmau534
🟧 hnOpenAI delayed its new model's development after the Hugging Face hacksbulaev42
🟠 redditOpenAI pulling back Astra
LocalLLaMA
OvertaxedOne440
🟠 redditAstra is about to drop 👀
OpenAI
Character_Novel_259220943
🟧 hnAjeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyurivish30
🟠 redditOpenAl's chief scientist on the neuralese controversy
singularity
Ok_Display_315921679
🟠 redditFable 5.1 outperforms Fable 5, Opus 5, and GPT‑5.6 Sol. Isn't time to reveal GPT-6 Astra?
OpenAI
RFOK24085

Interpretation history

Decision trace