2026-10-11 17:12 UTC

Independent use will determine whether SpecJudge can use CLAUDE.md, AGENTS.md, and nested repository instructions to select cheaper coding models without materially reducing task quality.

state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agents model-routing repository-contextSpecJudge

What is this?

SpecJudge is presented as a local tool that reads repository instruction files such as CLAUDE.md and AGENTS.md to recommend the least expensive coding model likely to handle that repository, while flagging insufficient context. The supplied results establish that these files encode repository-specific guidance, can be nested with closest-file precedence, and are used across multiple coding agents, but evidence on their value conflicts: some testing reports benefits, while other studies report higher token use and little or negative quality improvement. The snippets do not independently document SpecJudge’s creator, implementation, benchmark results, or whether instruction files contain enough signal for reliable model routing, so its core claim remains unverified here.

Why it matters to Scott

SpecJudge combines Scott’s existing CLAUDE.md-as-repository-spec pattern with his task-aware, cost-tiered model-routing work, while adding a concrete repository-level routing and insufficient-context decision. Independent results could affect how he builds routing into coding-agent harnesses—especially given his warning that bloated instruction files can increase cost and reduce success—but the tool’s claims and creator remain unverified, limiting relevance for now.
ip:concept.claude-md-patterndev:concept.task-aware-model-routingip:concept.fat-agents-md-anti-patternip:concept.evaluation-driven-developmentradar:concept.model-routingradar:multi-model-orchestrator-worker-agentsradar:karpathy-claude-md-rulesradar:concept.agent-evaluation
queries asked of Scott's wikis
  • repository-aware coding model routing
  • cheap-model escalation for coding agents
  • CLAUDE.md and AGENTS.md as machine-readable repository specifications
  • nested instruction precedence in agent harnesses
  • task quality measurement for coding-agent routing
  • insufficient-context detection and routing confidence

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Built a local tool that reads your CLAUDE.md and tells you which model that repo actually needs — and labels the answer when your context file isn't enough to tell
ClaudeAI
jokiruiz02
🟠 redditStop blindly using Claude Sonnet. I built an open-source evaluator (SpecJudge) to prove local models can beat it if you use Spec-Driven Development
LocalLLaMA
jokiruiz09
🟠 redditI built a tool that tells you whether your project actually needs Opus — it's now a one-command install in spec-kit
ClaudeAI
jokiruiz05

Interpretation history

Decision trace