2026-10-11 17:10 UTC

jabulari's measurement of 67,074 public OpenHands runs claims 77.8% of coding-agent runs carry at least one request with a stale post-edit file view (about 1 in 7 requests) because original reads persist after edits; adoption of state-tracking or auto-refresh mitigations would establish post-edit context staleness as a material harness failure mode.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highagent-harnesses context-staleness agent-evaluation

What is this?

A Reddit user 'jabulari' analyzed 67,074 public OpenHands (coding agent) runs and reports that 77.8% of runs contain at least one request where the agent's prompt includes a stale post-edit file view โ€” the original file read persists in context after the agent has edited that file. A second independent Reddit measurement (Sea-Perception1619, Claude Code handoff notes) found 16% of carried PR/issue state references were already wrong when read. Both posts have minimal engagement (score ~1, few comments) and no official vendor response or reproduction yet. The web search returned no external coverage; grounding rests entirely on the two Reddit self-reports.

Why it matters to Scott

Two independent Reddit measurements quantify a context-staleness failure mode (stale post-edit file views in OpenHands; stale PR/issue refs in Claude Code handoff notes) that Scott's canon already identifies as a first-class harness integrity problem: his context-engineering framework treats the context window as an attention budget requiring structural forgetting; signal-density is a first-class context-quality metric; stale-context-check is a validity gate for historical judgment; long-running-agents architecture mandates context hygiene and externalized durable state; model-plus-harness-benchmark-unit makes harness-induced failure a measured property; evaluation-driven-development gates releases on offline evaluation; agent-receipts demand reconstructable traces. The measurements give Scott dated, quantified evidence for a failure class his frameworks already name โ€” a publishing and audit opportunity.
dev:project.coding-agent-harnessdev:project.agent-loopdev:project.askip:framework.context-engineeringip:concept.signal-densityip:concept.stale-context-checkip:framework.long-running-agentsip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentip:concept.agent-receiptsip:concept.context-rotip:concept.context-bloatip:concept.attention-diffusiondev:concept.persistent-delta-event-logdev:concept.stale-while-revalidatedev:concept.agent-authored-context-compactionip:concept.checkpoint-disciplineip:concept.session-isolationip:concept.working-fidelityip:source.handover-notes-for-robots-ebookip:framework.three-tier-error-budgetsip:concept.answer-failure-classesip:source.observability-for-agentic-systems-what-to-log-how-to-redact-how-to-debug-ebookip:source.production-ready-ai-systems-ebookip:source.the-intelligent-rfp-ebookradar:concept.agent-harnessesradar:concept.agent-evaluationradar:concept.context-managementradar:concept.agent-benchmarksradar:agentgauntlet-failure-benchmarkradar:frontierharness-17x-cost-variationradar:harnessopt-agent-harness-optimization-benchmarkradar:ship-harness-benchradar:swe-bench-pro-harness-cost-parityradar:cache-hunter-prompt-cache-debuggingradar:ontoprune-context-pruningradar:futureos-context-compaction-recallradar:distil-decision-equivalent-context-compressionradar:headroom-reversible-context-compressionradar:recirculation-running-contextradar:coalent-source-aware-cache-invalidationradar:driftwatch-ast-token-pruningradar:automaton-durable-agent-stateradar:agent-memory-leaderboard-validationradar:concept.agent-memoryradar:concept.persistent-agentsradar:charter-durable-agent-control-planeradar:dropstone-persistent-agent-runtimeradar:headlong-persistent-agent-microharnessradar:exo-self-modifying-agent-harnessradar:seed-self-modifying-agent-harnessradar:evoharnessrl-self-evolving-agent-harnessradar:agent6-jailed-state-machine-harnessradar:hermes-agent-open-harnessradar:hermes-missions-durable-agent-executionradar:oh-my-subagents-persistent-workflowsradar:session-migrate-cross-harness-portabilityradar:teleport-cross-harness-session-portabilityradar:txcript-cross-harness-session-conversionradar:cache-tax-idle-session-warmingradar:awareness-local-agent-memoryradar:anansi-open-memory-apiradar:deposition-claude-code-memoryradar:drive9-agent-filesystem
queries asked of Scott's wikis
  • ip:concept.agent-harnesses evaluation failure modes
  • ip:concept.context-integrity prompt-assembly stale-file-view
  • dev:project.coding-agent-harness state-tracking auto-refresh
  • ip:concept.agent-evaluation measurement methodology context-staleness
  • ip:ebook.agent-memory file-state persistence across turns
  • dev:project.agent-loop file-view-cache invalidation

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p33momentum: steady1 platformsage 314h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-28 13:39โญ origin directly observedYour coding agent's prompt often contains old versions of files it already edited. Here's what I measured, and what I learned trying to fix it
jabulari on r/ClaudeAI
โ€”
10-10 03:15first on r/ClaudeAI ยท published ยท +277.6h49 of 307 PR and issue states in my Claude Code handoff notes were already wrong when a later session read them
Sea-Perception1619
โ€”
09-28 13:39amplified on r/ClaudeAIreddit.post.1wsewkf
jabulari
peak 1 ยท 1 comments ยท 15% of case engagement
10-10 03:15amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1x24pot
Sea-Perception1619
peak 1 ยท 10 comments ยท 84% of case engagement
09-28 15:20our radar first saw it ยท +1.7hdiscovery anchor: reddit.post.1wsewkfโ€”
pace: p22 vs 1188 stories at the 168h mark (now 314h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญYour coding agent's prompt often contains old versions of files it already edited. Here's what I measured, and what I learned trying to fix it
ClaudeAI
jabulari11
๐ŸŸ  reddit49 of 307 PR and issue states in my Claude Code handoff notes were already wrong when a later session read them
ClaudeAI
Sea-Perception1619010

Interpretation history

Decision trace