2026-10-11 16:37 UTC

Reddit builder maverick_man1111's code-level audit claims 13 cases where seven widely used LLM-eval tools (NVIDIA SkillEvaluator, the agent-skills harness, MLflow, LangSmith, DSPy, DeepEval, Harbor) return scores not backed by what they measure โ€” three in his own plugin โ€” and maintainer fixes plus third-party replication decide whether eval-score validity becomes a recognized, tracked gap in agent evaluation.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation eval-tooling benchmark-validitymaverick_man1111

What is this?

A Reddit builder, maverick_man1111, posted a code-level audit claiming 13 cases across seven LLM/agent evaluation tools (NVIDIA SkillEvaluator, the agent-skills harness, MLflow, LangSmith, DSPy, DeepEval, Harbor) where returned scores are not backed by what the tool actually measures โ€” including three in his own Claude Code plugin. The supplied web results confirm the tool landscape (MLflow, DeepEval, LangSmith and peers are heavily deployed, and DeepEval is explicitly positioning itself as the eval harness for Claude Code/Cursor-style coding agents) and show LLM-judge validity is an active discourse topic โ€” MachineLearningMastery warns that judge position/self-preference/verbosity biases 'show up... inside the very frameworks' it reviews. However, none of the snippets surface the Reddit audit itself, the NVIDIA or Harbor tools, or any maintainer response, so the audit's specific claims, its repros, and any fixes remain unverified by this evidence.

Why it matters to Scott

maverick_man1111 โ€” already tracked on this radar as Driftproof's creator โ€” independently lands on Scott's measured-vs-claimed position with code-level, nameable receipts against the standard eval harnesses (LangSmith, MLflow, DSPy, DeepEval โ€” the exact toolset his own wikis' matched queries on evaluation-driven projects surface), and disclosing three defects in his own plugin is Challenger-Never-Arbiter performed in the wild: an external challenger converting objection into instrumentation. Not a repetition of canon โ€” his wikis hold the principle (perceptual coverage first on any judge, independent key audits, evidence spent at its class, measures as the arbiter), and this case supplies the named-tool inventory and the maintainer-fix/replication test that turns 'verify the verifier' into a concrete tool-selection question for the very gates evaluation-driven development ships through, plus a dated-receipts publish.
ip:concept.evaluation-driven-developmentip:framework.challenger-never-arbiterip:concept.perceptual-coverageip:concept.evidence-class-ladderdev:concept.llm-rubric-gradingdev:concept.review-until-clear-loopradar:driftproof-skill-regression-testingradar:anthropic-claude-eval-pluginradar:agent-skill-bloat-gradingradar:ai-benchmark-saturation-distortionradar:llm-judge-prior-score-anchoring
queries asked of Scott's wikis
  • LLM-as-judge validity bias trusting scores
  • agent evaluation harness scoring design
  • Claude Code skills grading audit
  • benchmark validity measured vs claimed
  • eval-to-patch loop coding agent
  • MLflow LangSmith DSPy DeepEval in my projects

Measured heat

now 0 pts/hpeak 9 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 290h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-29 14:00โญ origin echo-reconstructedPreprint "Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools" (Version 1, 30 September 2026
driftproofhq on paper (echo) ยท attributed from reddit.post.1wywvmy
โ€”
10-06 08:01first on r/ClaudeAI ยท published ยท +162.0hI read the scoring code of 7 tools people use to grade Claude Code skills and LLM apps (NVIDIA, MLflow, LangSmith, DSPy, DeepEval...). Found 13 cases where the result wasn't backed by what the tool actually measured. 3 were in my own Claude Code plugin.
maverick_man1111
โ€”
10-06 08:01amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wywvmy
maverick_man1111
peak 3 ยท 1 comments ยท 102% of case engagement
10-06 08:20our radar first saw it ยท +162.3hdiscovery anchor: reddit.post.1wywvmyโ€”
pace: p35 vs 1188 stories at the 168h mark (now 290h old) โ€” ahead of agentgate-signed-agent-receipts (1.3x), behind agentic-determinism-index (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditI read the scoring code of 7 tools people use to grade Claude Code skills and LLM apps (NVIDIA, MLflow, LangSmith, DSPy, DeepEval...). Found 13 cases where the result wasn't backed by what the tool actually measured. 3 were in my own Claude Code plugin.
ClaudeAI
maverick_man111131
๐ŸŸง echo.paper โญPreprint "Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools" (Version 1, 30 September 2026driftproofhqโ€”โ€”

Interpretation history

Decision trace