2026-10-11 17:12 UTC

Independent training runs will determine whether Financial-RLVR-10K’s execution-verified rewards produce useful financial-reasoning gains without relying on LLM-based reward judges.

state: expiredheat: lowuncertainty: highknownscott: lowrlvr open-model-training verifiable-rewardscoslinedev

What is this?

Financial-RLVR-10K is presented as a dataset from coslinedev containing 10,000 financial-reasoning problems whose answers are checked through sandbox execution rather than an LLM reward judge. Its README reportedly claims all records are sandbox-verified and that 1,950 examples contain adversarial financial logic traps, making it intended for GRPO/RLVR training with programmatic rewards. The supplied web snippets explain the broader RLVR method—rewarding outputs that pass rule-based checks—but provide no independent training results or direct corroboration of the dataset’s quality, verification coverage, or claimed gains.

Why it matters to Scott

Scott already holds the core position in “Mechanically Different Verifiers” and “Correlated Checkers Pitfall”: independent, ground-truth checks are epistemically stronger than model-judge agreement. This release is a potentially useful RLVR implementation of that position, but without independent training results it is currently another example rather than evidence that would change what he builds or argues.
ip:concept.mechanically-different-verifiersip:concept.correlated-checkers-pitfallip:concept.specification-gamingip:concept.evaluation-driven-developmentdev:concept.synthetic-finetuning-datasetradar:concept.reinforcement-learningradar:concept.verificationradar:concept.llm-evaluationradar:nanorl-lightweight-llm-rl-trainer
queries asked of Scott's wikis
  • execution-verified rewards versus LLM judges
  • RLVR and GRPO training strategy
  • verifier design and reward hacking
  • sandboxed code execution for model training
  • financial reasoning benchmarks and adversarial traps
  • open-model post-training datasets

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Dataset Release] I built 10k execution-verified financial problems (with 1,950 logic traps) so you don't need LLM-as-a-judge for GRPO
LocalLLaMA
coslinedev83
🟧 echo.other ⭐The dataset README says: "10,000 verified records," "100% Sandbox-Verified," and "1,950 samples (19.5%)" with adversarial financial traps. Icoslinedev——

Interpretation history

Decision trace