Driftproof creator maverick_man1111 claims its released tool compares scored runs with and without agent instructions across selected models and preserves dated, hashed records, enabling detection of instruction regressions after model or configuration changes.
state: seedheat: lowuncertainty: mediumknownscott: lowagent-harnesses coding-agents evaluationmaverick_man1111Driftproof
What is this?
The case describes Driftproof as a released agent-instruction evaluation tool whose claimed creator, maverick_man1111, says it compares scored runs with and without instructions across selected models and retains dated, hashed records to detect regressions. The supplied evidence title additionally claims that two launch-day auditors broke three things, but provides no details of those failures or fixes. None of the web snippets directly identifies Driftproof or its creator; they cover other drift-monitoring and agent-run tooling, so they do not independently establish the release, implementation, or effectiveness of the claimed regression detection.
Why it matters to Scott
Driftproof’s claimed instruction-ablation tests and retained run records repeat positions already held in Scott’s Reflexive Agent Design, Evaluation-Driven Development, and Trace-backed agent comparison pages. No supplied radar hit tracks Driftproof itself, but the creator-only account establishes neither a consequential new adopter nor a validated capability that would change Scott’s evaluation practice; it remains a potential tool to inspect rather than a demonstrated advance.
ip:framework.reflexive-agent-designip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisonradar:concept.agent-evaluationradar:ship-harness-benchradar:agent-skill-bloat-gradingradar:ai-stupid-level-benchmark-drift
queries asked of Scott's wikis
- agent harness instruction ablation baseline evaluations
- coding agent skills regression tests model upgrades
- prompt effectiveness cross-model configuration testing
- reproducible agent runs hashed artifacts evaluation provenance
- evaluation scoring reliability stochastic agent behavior
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 671h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p32 vs 1032 stories at the 336h mark (now 671h old) — ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-13T17:32:24Z
grounded: known/low — Driftproof’s claimed instruction-ablation tests and retained run records repeat positions already held in Scott’s Reflexive Agent Design, Evaluation-Driven Deve
2026-09-13T17:28:07Z
case created — The release describes a specific longitudinal regression workflow distinct from existing skill-grading and provenance cases.
Decision trace
- 09-14 03:32groundDriftproof’s claimed instruction-ablation tests and retained run records repeat positions already held in Scott’s Reflexive Agent Design, Evaluation-Driven Development, and Trace-backed agent comparis
- 09-14 03:28createThe release describes a specific longitudinal regression workflow distinct from existing skill-grading and provenance cases.