Coding Atlas’s publisher claims to have released every diff and transcript from coding agents operating on six booby-trapped repositories, potentially making hostile-repository behavior directly auditable.
state: seedheat: lowuncertainty: highknownscott: lowcoding-agents agentic-security model-evaluationtap2k
What is this?
The case describes Coding Atlas as a release of coding-agent runs on six booby-trapped repositories, with its publisher claiming that every diff and transcript is public; it names tap2k but supplies no verified publisher identity or role. None of the supplied web snippets corroborates that release, its completeness, or the hostile-repository experiments. The GitHub result for pacifio/atlas describes source control linking agent commits to session prompts, tool calls, and reasoning, but the supplied material does not establish that it is Coding Atlas or connected to this experiment.
Why it matters to Scott
Publishing agent paths for inspection repeats Scott’s existing position in Reflexive Agent Design and Trace-backed agent comparison; the supplied evidence establishes neither a consequential new adopter nor findings that would change his harness design. Verified hostile-repository traces could become useful test material, but this release and its completeness remain uncorroborated; the radar tracks related repository-injection testing, not this specific Coding Atlas development.
ip:framework.reflexive-agent-designdev:concept.trace-backed-agent-comparisonradar:repository-content-agent-injectionradar:concept.coding-agent-securityradar:concept.agent-evaluation
queries asked of Scott's wikis
- coding agent hostile repository trust boundaries
- repository prompt injection agent harness security
- agent evaluation reproducible traces diffs transcripts
- agent action provenance audit trails
- adversarial coding benchmarks harness failure analysis
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 745h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p11 vs 519 stories at the 720h mark (now 745h old) — behind addom-local-coding-harness (0.5x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-10T16:41:17Z
No substantive new evidence changes the interpretation: the GitHub echo repeats the submission’s claim rather than independently establishing accessible, complete evaluation artifacts. This remains potentially useful hostile-repository test material, without findings or reusable details that bear on Scott’s harness decisions.
2026-09-10T15:53:05Z
grounded: known/low — Publishing agent paths for inspection repeats Scott’s existing position in Reflexive Agent Design and Trace-backed agent comparison; the supplied evidence estab
2026-09-10T15:50:56Z
case created — This is a distinct published evaluation artifact; the evidence does not connect it to the existing supply-chain-trust paper or establish its findings.
Decision trace
- 10-04 04:56review_dormantscheduled targets exhausted or 28 quiet days
- 10-04 04:56drop_targetsquiet through full ladder or over cap 8
- 09-11 02:41repriceNo substantive new evidence changes the interpretation: the GitHub echo repeats the submission’s claim rather than independently establishing accessible, complete evaluation artifacts. This remains po
- 09-11 02:41alert_silentThere is no new consequential delta. The announced publication can wait for a briefing because the available material identifies neither concrete agent failures nor actionable test details; the reason
- 09-11 02:41alert_routeThere is no new consequential delta. The announced publication can wait for a briefing because the available material identifies neither concrete agent failures nor actionable test details; the reason
- 09-11 01:54alert_silentThe publisher’s announcement introduces potentially useful hostile-repository test material, but the supplied evidence contains no concrete failure findings, affected agents, or reusable test details
- 09-11 01:54surface_candidateThe publisher’s announcement introduces potentially useful hostile-repository test material, but the supplied evidence contains no concrete failure findings, affected agents, or reusable test details
- 09-11 01:54alert_routeThe publisher’s announcement introduces potentially useful hostile-repository test material, but the supplied evidence contains no concrete failure findings, affected agents, or reusable test details
- 09-11 01:53groundPublishing agent paths for inspection repeats Scott’s existing position in Reflexive Agent Design and Trace-backed agent comparison; the supplied evidence establishes neither a consequential new adopt
- 09-11 01:50createThis is a distinct published evaluation artifact; the evidence does not connect it to the existing supply-chain-trust paper or establish its findings.