2026-10-11 17:10 UTC

Pairmark’s maintainer claims its released isolated-worktree harness, automated checks, and blind reciprocal patch reviews provide a practical per-repository method for comparing Claude Code and Codex on real tasks.

state: expiredheat: lowuncertainty: highknownscott: mediumcoding-agents agent-harnesses agent-evaluationHemanshu UpadhyayPairmarkAnthropicOpenAI

What is this?

Pairmark is presented as a released repository-level harness maintained by Hemanshu Upadhyay for comparing Claude Code and Codex on the same coding task. According to the evidence titles, it gives each agent an isolated worktree, runs automated checks, and uses blind reciprocal review of their patches; the maintainer reports one trial in which Codex fixed a build-script bug but failed a rule. The supplied web snippets discuss broader Claude Code–Codex comparisons and cross-review workflows but do not independently establish Pairmark’s implementation, methodology, or results, so those details remain maintainer claims.

Why it matters to Scott

Scott already holds and implements this methodology in “Trace-backed agent comparison” and “Model-Plus-Harness Benchmark Unit,” while the radar already tracks repository-level coding-agent benchmarks and cross-model code-review validation. Pairmark is therefore not a new position, but its released per-repository harness is directly testable against Scott’s active evaluation work; its reciprocal judging also raises his existing concern that model judges are not independent unless backed by mechanically different checks.
dev:concept.trace-backed-agent-comparisonip:concept.model-plus-harness-benchmark-unitip:concept.mechanically-different-verifiersip:concept.evaluation-driven-developmentradar:cross-model-code-review-validationradar:concept.coding-agent-benchmarksradar:ai-to-ai-pr-review
queries asked of Scott's wikis
  • coding-agent evaluation on real repositories
  • isolated worktrees for parallel coding agents
  • blind cross-model review of agent patches
  • automated checks as agent evaluation harness
  • reciprocal judging bias between LLM agents
  • repository-specific benchmarks versus standardized coding benchmarks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI raced Claude Code against Codex on the same task with blind cross-judging. Codex fixed a bug in my build script and lost on a rule.
ClaudeAI
Infamous_Term_96522
🟧 hnShow HN: Pairmark, race Claude Code vs. Codex on your repo, blind cross-judgedhemanshu41210
🟧 echo.github ⭐The Pairmark repository releases a harness for racing Claude Code and Codex on the same repository task with isolated worktrees, automated cHemanshu Upadhyay——

Interpretation history

Decision trace