TaxCalcBench is a benchmark for evaluating whether frontier models can calculate simplified US personal income-tax returns; its paper reports that state-of-the-art models solved fewer than one-third of federal cases. The case concerns OpenTax, described in the supplied evidence titles as an open-source, verifiable deterministic tax engine that purportedly raises a Sonnet agent from about 6% to 96% exact-return accuracy. However, the provided web snippets do not establish OpenTax’s affiliation with Invaro AI, independently reproduce that gain, or confirm that the remaining two failures are benchmark errors; those claims remain unverified here.
No intersection found: the supplied Scott wiki and radar searches returned no hits, so there is no grounded basis to connect OpenTax’s unverified benchmark claims to a position, project, or tracked story.
queries asked of Scott's wikis
- deterministic tools versus LLM reasoning
- tool-augmented agents and exact computation
- MCP deterministic tool integration
- verifiable computation for high-stakes AI
- benchmark error auditing and independent reproduction
- AI tax and compliance systems
2026-08-07T17:41:52Z
The initial validation window produced no independent reproduction or direct benchmark-maintainer evidence, and no substantive delta has appeared after repeated amplification. The claim remains unresolved rather than disproved, but it no longer warrants an active case unless external validation emerges.
2026-07-25T01:21:39Z
The newly attached material adds no independent reproduction or direct benchmark-maintainer evidence; the headline result still rests entirely on OpenTax’s own claim. Repetitive amplification is exhausted, so the case should stay dormant until substantive external validation appears.
2026-07-24T16:29:10Z
The attachment provides no new independent line; the benchmark gain and alleged test-case errors still trace solely to OpenTax. Repetitive amplification is exhausted, so the case should remain dormant until an external reproduction or direct maintainer statement appears.
2026-07-24T14:31:25Z
No substantive new line of evidence is present: the benchmark gain and claimed benchmark errors still trace to OpenTax rather than an independent reproduction or direct maintainer statement. Further engagement-only updates should not revive the case; wait for external validation.
2026-07-24T13:25:47Z
The attachment adds no independent reproduction or direct maintainer evidence; the case remains entirely dependent on OpenTax’s originating claims. Repetitive amplification is exhausted, so review should pause until substantive external validation appears.
2026-07-24T12:28:14Z
The newly attached evidence remains another observation of OpenTax’s originating claim, not an independent reproduction or direct benchmark-maintainer statement. The validation question remains open, but repetitive amplification adds no meaning and warrants a slower cadence.
2026-07-24T11:27:56Z
The attachment adds no independent evidence; the benchmark result and alleged maintainer confirmation still trace back to OpenTax. Repeated amplification is no longer informative, so the case should remain dormant pending an external reproduction or direct maintainer statement.
2026-07-24T10:28:47Z
The newly attached evidence adds no independent line: all claims still trace to OpenTax, including the alleged maintainer confirmation. Repetitive amplification no longer merits frequent review; wait for a reproducible implementation report or direct benchmark-maintainer evidence.
2026-07-24T08:26:39Z
The evidence remains circular amplification of OpenTax’s own claim, with no independent reproduction or direct benchmark-maintainer confirmation. Repeated engagement updates no longer justify hourly review; the case should wait for substantive external validation.
2026-07-24T07:28:51Z
The attached evidence still originates with OpenTax and adds neither an independent reproduction nor direct benchmark-maintainer confirmation. Repeated amplification has not changed the case’s meaning, so it remains a low-attention validation watch.
2026-07-24T06:26:53Z
The newly attached material still traces back to OpenTax’s own benchmark claim and adds no independent reproduction or maintainer confirmation. The case remains testable but unchanged in meaning, with continued amplification insufficient for corroboration.
2026-07-24T05:27:03Z
Still no independent reproduction or benchmark-maintainer confirmation; evidence set unchanged in substance since last look, just re-flagged as dirty. Amplification without corroboration, relevance remains none.
2026-07-24T04:21:00Z
No independent reproduction or benchmark-maintainer confirmation has emerged; the attached observations remain amplification of OpenTax’s own claim. With engagement flat to declining, the case still hinges entirely on external validation.
2026-07-24T03:25:32Z
grounded: novel/none — No intersection found: the supplied Scott wiki and radar searches returned no hits, so there is no grounded basis to connect OpenTax’s unverified benchmark clai
2026-07-24T03:24:34Z
origin walked (codex/luna, conf 0.95): anchor hn.story.49030618 -> echo.github.fc1de5a4e9 by Asmi Gulati
2026-07-24T03:22:32Z
case created — Two cross-platform observations point to a concrete tool and unusually large, independently testable benchmark improvement.