Independent reproduction and OpenAI’s response will determine whether a Codex update around July 22 introduced persistent agent loops that materially increase token usage on otherwise achievable tasks.
state: expiredheat: lowuncertainty: highknownscott: mediumcodex coding-agents agent-harnessesOpenAI
What is this?
A user report titled 'Something is wrong with Codex since July 22' claims a Codex update around that date introduced persistent agent-loop behavior that increases token usage on tasks that were previously achievable more efficiently. The web results confirm OpenAI shipped GPT-5-Codex upgrades (announced July 23) emphasizing 'persistent, independent execution' on longer tasks and dynamic thinking-time allocation, which is consistent with the mechanism the complaint points at, but there is no independent reproduction or OpenAI acknowledgment of a regression in the supplied material—only the original complaint and OpenAI's own framing of the feature as an improvement.
Why it matters to Scott
Known through the 12-Factor Agents Framework’s Observable Autonomy position and the Ask terminal agent, which actively uses Codex in a recursive tool loop. If independently reproduced, the regression would directly affect Scott’s loop controls, completion criteria, model routing, and token economics; for now it remains a single unconfirmed complaint rather than evidence requiring a changed position.
ip:framework.12-factor-agents-frameworkdev:project.askdev:project.llmreportradar:concept.long-horizon-agentsradar:concept.agent-harnessesradar:concept.coding-agentsradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
- agent harness loop control and runaway execution
- coding agent token economics and cost efficiency tradeoffs
- agent autonomy vs interruptibility design principles
- Codex CLI or coding agent tooling notes
- persistent execution vs task completion criteria in agent design
- known failure modes of long-running autonomous agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-07T21:34:15Z
The refreshed discussion still provides no independent reproduction; comments instead add skepticism, adjacent complaints, and calls for a fixed benchmark. The version-bounded regression report has not developed within its useful window and no confirming event is presently expected.
2026-08-05T21:24:58Z
The newly attached material adds no independent reproduction or OpenAI acknowledgment; the case remains a single low-engagement report. Without corroboration, the apparent regression is still indistinguishable from a workload-, configuration-, or project-specific failure mode.
2026-08-05T11:25:33Z
grounded: known/medium — Known through the 12-Factor Agents Framework’s Observable Autonomy position and the Ask terminal agent, which actively uses Codex in a recursive tool loop. If i
2026-08-05T11:21:52Z
case created — A heavy-use team reports a version-bounded reliability and cost regression that is distinct from existing Codex cases and merits targeted reproduction.
Decision trace
- 08-08 07:34expireThe refreshed discussion still provides no independent reproduction; comments instead add skepticism, adjacent complaints, and calls for a fixed benchmark. The version-bounded regression report has no
- 08-08 07:34alert_silentThe new comments neither confirm the reported looping regression nor establish an OpenAI response, so they do not justify interrupting Scott or carrying the episode forward.
- 08-08 07:34alert_routeThe new comments neither confirm the reported looping regression nor establish an OpenAI response, so they do not justify interrupting Scott or carrying the episode forward.
- 08-06 07:24repriceThe newly attached material adds no independent reproduction or OpenAI acknowledgment; the case remains a single low-engagement report. Without corroboration, the apparent regression is still indistin
- 08-06 07:20mark_dirtyengagement_update
- 08-05 21:25groundKnown through the 12-Factor Agents Framework’s Observable Autonomy position and the Ask terminal agent, which actively uses Codex in a recursive tool loop. If independently reproduced, the regression
- 08-05 21:21createA heavy-use team reports a version-bounded reliability and cost regression that is distinct from existing Codex cases and merits targeted reproduction.