2026-10-11 17:11 UTC

Independent use will determine whether Claude Code System Prompts Time Machine accurately archives versioned system prompts and enables reproducible evaluation of coding-agent harness changes.

state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agents agent-harnesses prompt-evaluationbrewpirateAnthropic

What is this?

Claude Code System Prompts Time Machine is described as an archive and extraction tool for tracking Claude Code’s system prompts across CLI/SDK releases, associated with the Piebald-AI repository; the supplied snippets do not establish brewpirate’s precise role. The repository says it extracts prompts directly from published Claude Code npm code and maintains prompt/token data plus a changelog spanning 253 versions, while a session-level evaluation workflow is reportedly still under development. The surrounding evidence supports the motivation—small harness-prompt changes can materially affect agent behavior—but independent accuracy and reproducibility testing are not yet established here.

Why it matters to Scott

The tool independently operationalizes Scott’s position that prompts and harnesses are versioned production artifacts whose behavioral effects should be tested against repeatable, trace-backed fixtures—not attributed to the model alone. It could become useful infrastructure for his agent-comparison and session-archive work, but accuracy is unverified and the evaluation workflow remains under development.
ip:concept.evaluation-driven-developmentip:framework.the-prompt-is-sourceip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisondev:project.search-conversationsradar:concept.agent-harnessesradar:concept.agent-evaluationradar:stencil-harness-coding-improvement
queries asked of Scott's wikis
  • coding-agent prompt and harness versioning
  • reproducible evaluation of agent harness changes
  • detecting behavioral drift across coding-agent releases
  • system-prompt provenance and artifact capture
  • session replay and regression testing for agents
  • model-versus-harness performance attribution

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditClaude Code System Prompts 1.0.0 - 2.1.232 cli/sdk
ClaudeAI
dopamine_91111
🟧 echo.github ⭐An archive and capture tool for generated Claude Code system prompts, with a session-level evaluation workflow under development.brewpirate——
🟧 hnClaude: System Promptstosh756274

Interpretation history

Decision trace