2026-10-11 16:38 UTC

tamewild claims its released PCSS recipe fine-tunes Qwen3-4B Base on 100 zebra puzzles in roughly 6.5 GPU minutes to reach 85.26% on the full MATH benchmark, a reported 31.16-point gain that could make substantial small-model reasoning improvements unusually cheap.

state: seedheat: lowuncertainty: highconvergesscott: lowsmall-model-training fine-tuning math-benchmarkstamewild

What is this?

The case attributes to tamewild a released fine-tuning recipe called PCSS, claiming that training Qwen3-4B Base on 100 zebra puzzles for roughly 6.5 GPU minutes produces 85.26% accuracy on full MATH, a 31.16-percentage-point gain. The supplied web snippets establish Qwen3-4B-Base as a four-billion-parameter model from the Qwen Team, but none corroborates tamewild’s identity, PCSS, the notebook release, training time, or reported results; the other fine-tuned models returned are separate projects. The benchmark scope also remains unresolved: the evidence title names MATH-500, whereas the hypothesis names full MATH.

Why it matters to Scott

At the approach level, the claimed synthetic-puzzle fine-tuning converges with Scott’s Synthetic fine-tuning dataset generation concept and Salesforce fine-tuning-data factory, but does not establish useful transfer to his domain-training workflow. The unusually cheap reasoning gain remains uncorroborated, with MATH versus MATH-500 unresolved, so this is presently another synthetic-training example rather than grounds to change his builds; the related radar episodes do not track this same development.
dev:concept.synthetic-finetuning-datasetdev:project.redditradar:500-dollar-9b-rl-catalog-reviewradar:scaffold-cot-small-model-datasetradar:concept.fine-tuning
queries asked of Scott's wikis
  • small-model fine-tuning economics versus frontier APIs
  • reasoning transfer from synthetic logic puzzles
  • training data quality versus quantity
  • benchmark validity reproducible evaluation harnesses
  • local model customization and reasoning capability

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 746h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-10 14:00⭐ origin echo-reconstructedReports mathematical reasoning gains after fine-tuning small base models on 100–500 zebra-puzzle traces without mathematical training exampl
tamewild on blog (echo) Β· attributed from reddit.post.1wdhb24
β€”
09-11 14:06first on r/LocalLLaMA Β· published Β· +24.1hFine-tuning Qwen 3 4B Base on 100 zebra puzzles yielded +31% on MATH-500. 6.5-min (Single H100/H200) reproduction notebook included.
TGSCrust
β€”
09-30 17:12first on hacker news Β· published Β· +483.2hCan My Puzzle Write-Ups 'Lift' Cheap LLMs?
speckx
β€”
09-11 14:06amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wdhb24
TGSCrust
peak 37 Β· 7 comments Β· 96% of case engagement
09-30 17:12amplified on hacker newshn.story.49911694
speckx
peak 1 Β· 0 comments Β· 4% of case engagement
09-11 14:20our radar first saw it Β· +24.3hdiscovery anchor: reddit.post.1wdhb24β€”
pace: p62 vs 519 stories at the 720h mark (now 746h old) β€” ahead of openai-bio-bug-bounty (1.0x), behind local-kv-cache-pressure-probe (0.9x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditFine-tuning Qwen 3 4B Base on 100 zebra puzzles yielded +31% on MATH-500. 6.5-min (Single H100/H200) reproduction notebook included.
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-11T14:22:15.938573+00:00

Back to Articles Eliciting Reasoning with 100-500 Zebra Puzzles Community Article Published
					September 11, 2026 Upvote 1 tamewild tamewild Follow Resources & Quick Links: Read this Post: HuggingFace (Rendered Blog) | GitHub (Markdown Source) GitHub Repository: tamewild/pcss Reproduction Notebooks (Single H100/H200): Qwen 3 4B (~6.5 mins) | Granite 4.1 3B (~23 mins) | Qwen 3.5 9B (~40 mins) Models: PCSS-Qwen3-4B | PCSS-Granite-4.1-3B | PCSS-Qwen3.5-9B Datasets: zebra_100 | instruct5 1. Introduction Encountering the research on the Hyperfitting Phenomenon sparked an interest in exploring how pre-trained base models behave when micro-datasets are fine-tuned to near-zero training loss. We focused on training small pre-trained base models on 100 to 500 pure logic deduction traces (specifically 5x5 zebra puzzles) with zero mathematical data in the fine-tuning dataset . This synthetic dataset was generated by Qwen3-235B-A22B-Instruct-2507 in September 2025 using rejection sampling. Across multiple model families and architectures: Qwen 3 4B Base fine-tuned on 100 zebra puzzles (~6.5 minutes on 1 GPU) reaches 85.26% on the full 5,000-problem MATH benchmark (+31.16% delta over base) and 21.67% on AIME 2025 , rivaling the official post-trained model in non-thinking mode on MATH-500 while outperforming it on AIME 2025. Granite 4.1 3B Base fine-tuned on 500 zebra puzzles achieves 77.73% on MATH-500 and 19.44% on AIME 2025 , outperforming the official IBM Granite 4.1 3B Instruct model (+11.13 and +12.77 percentage points, respectively). Qwen 3.5 9B Base fine-tuned on 500 zebra puzzles jumps from 17.00% to 80.50% on 5x5 zebra puzzles (Pass@3: 97.05% ), while jumping from 5.42% to 22.92% on the uncontaminated ArXivMath 05/26 benchmark (Pass@3: 34.88% ) and reaching 96.60% on MATH-500 and 60.67% on AIME 2025 . While this out-of-domain generalization appears primarily driven by the pure logic dataset, we chose to manage the optimization process through adaptive loss scaling. For our primary results, we utilized PCSS (Per-Example Calibrated Sigmoid Scaler) : an adaptive, data-aware loss scaler derived directly from KTO. 2. The Mathematics of PCSS 2.1 From Preference Optimization to a Dynamic SFT Scaler PCSS originates from KTO . But why adapt an alignment algorithm for Supervised Fine-Tuning? The inspiration came from an experiment in the original KTO paper (Table 4). The authors demonstrated that running KTO using only desirable/positive examples (the exact same unpaired data used for SFT) yielded results highly competitive with standard SFT. While the paper conservatively notes this setup "yields similar results to SFT when Ξ² \beta Ξ² is small," the results with Ξ² = 0.1 \beta=0.1 Ξ² = 0.1 raised the overall average benchmark score from 29.7% to 31.1%, driven by a substantial relative increase on the GSM8K mathematical reasoning benchmark (from 1.0% to 12.5%). Why does an alignment algorithm behave this way? We can answer this by performing a gradient analysis on standard KTO when applied to purely desirable data. Setting the gain sensitivity to 1 ( Ξ» D = 1 ) (\lambda_D = 1) ( Ξ» D ​ = 1 ) , the KTO loss for a desirable ground-truth sequence y βˆ— y^* y βˆ— is: L KTO ( ΞΈ ) = 1 βˆ’ Οƒ ( Ξ² ( r ΞΈ ( x , y βˆ— ) βˆ’ z 0 ) ) \mathcal{L}_{\text{KTO}}(\theta) = 1 - \sigma(\beta(r_\theta(x, y^*) - z_0)) L KTO ​ ( ΞΈ ) = 1 βˆ’ Οƒ ( Ξ² ( r ΞΈ ​ ( x , y βˆ— ) βˆ’ z 0 ​ )) where r ΞΈ ( x , y ) r_\theta(x, y) r ΞΈ ​ ( x , y ) is the unnormalized log ratio ( log ⁑ Ο€ ΞΈ ( y ∣ x ) βˆ’ log ⁑ Ο€ ref ( y ∣ x ) ) (\log \pi_\theta(y|x) - \log \pi_{\text{ref}}(y|x)) ( lo g Ο€ ΞΈ ​ ( y ∣ x ) βˆ’ lo g Ο€ ref ​ ( y ∣ x )) . The loss for the ground-truth sequence is evaluated at r ΞΈ ( x , y βˆ— ) r_\theta(x, y^*) r ΞΈ ​ ( x , y βˆ— ) , and z 0 z_0 z 0 ​ is the KL divergence estimate. Crucially, KTO explicitly does not backpropagate through z 0 z_0 z 0 ​ . Taking the gradient with respect to ΞΈ \theta ΞΈ : βˆ‡ ΞΈ L KTO ( ΞΈ ) = βˆ’ Οƒ β€² ( Ξ² ( r ΞΈ ( x , y βˆ— ) βˆ’ z 0 ) ) β‹… Ξ² β‹… βˆ‡ ΞΈ r ΞΈ ( x , y βˆ— ) \nabla_\theta \mathcal{L}_{\text{KTO}}(\theta) = -\sigma'(\beta(r_\theta(x, y^*) - z_0)) \cdot \beta \cdot \nabla_\theta r_\theta(x, y^*) βˆ‡ ΞΈ ​ L KTO ​ ( ΞΈ ) = βˆ’ Οƒ β€² ( Ξ² ( r ΞΈ ​ ( x , y βˆ— ) βˆ’ z 0 ​ )) β‹… Ξ² β‹… βˆ‡ ΞΈ ​ r ΞΈ ​ ( x , y βˆ— ) Because the reference model is frozen, βˆ‡ ΞΈ r ΞΈ ( x , y βˆ— ) = βˆ‡ ΞΈ log ⁑ Ο€ ΞΈ ( y βˆ— ∣ x ) \nabla_\theta r_\theta(x, y^*) = \nabla_\theta \log \pi_\theta(y^*|x) βˆ‡ ΞΈ ​ r ΞΈ ​ ( x , y βˆ— ) = βˆ‡ ΞΈ ​ lo g Ο€ ΞΈ ​ ( y βˆ— ∣ x ) . Since the standard SFT loss L sft ( ΞΈ ) L_{\text{sft}}(\theta) L sft ​ ( ΞΈ ) is the negative mean log-probability over N N N tokens, log ⁑ Ο€ ΞΈ ( y βˆ— ∣ x ) = βˆ’ N β‹… L sft ( ΞΈ ) \log \pi_\theta(y^*|x) = -N \cdot L_{\text{sft}}(\theta) lo g Ο€ ΞΈ ​ ( y βˆ— ∣ x ) = βˆ’ N β‹… L sft ​ ( ΞΈ ) . Therefore, βˆ‡ ΞΈ log ⁑ Ο€ ΞΈ ( y βˆ— ∣ x ) = βˆ’ N β‹… βˆ‡ ΞΈ L sft ( ΞΈ ) \nabla_\theta \log \pi_\theta(y^*|x) = -N \cdot \nabla_\theta L_{\text{sft}}(\theta) βˆ‡ ΞΈ ​ lo g Ο€ ΞΈ ​ ( y βˆ— ∣ x ) = βˆ’ N β‹… βˆ‡ ΞΈ ​ L sft ​ ( ΞΈ ) . Substituting this back in gives the exact gradient update for standard KTO on positive data: βˆ‡ ΞΈ L KTO ( ΞΈ ) = [ Ξ² β‹… Οƒ β€² ( Ξ² ( r ΞΈ ( x , y βˆ— ) βˆ’ z 0 ) ) β‹… N ] βˆ‡ ΞΈ L sft ( ΞΈ ) \nabla_\theta \mathcal{L}_{\text{KTO}}(\theta) = \Big[ \beta \cdot \sigma'(\beta(r_\theta(x, y^*) - z_0)) \cdot N \Big] \nabla_\theta L_{\text{sft}}(\theta) βˆ‡ ΞΈ ​ L KTO ​ ( ΞΈ ) = [ Ξ² β‹… Οƒ β€² ( Ξ² ( r ΞΈ ​ ( x , y βˆ— ) βˆ’ z 0 ​ )) β‹… N ] βˆ‡ ΞΈ ​ L sft ​ ( ΞΈ ) When you apply KTO to purely desirable examples, it is no longer performing preference alignment against negative data. Instead, it mathematically reduces to the standard SFT gradient multiplied by a dynamic scaler . However, this same equation highlights why standard KTO would struggle on logic chains spanning up to 11,000+ tokens: Length-Induced Gradient Inflation: The scaler is multiplied by the sequence length N N N . For long traces, the gradient magnitude increases sharply by a factor of hundreds. Length Bias & Gradient Saturation: Because r ΞΈ ( x , y βˆ— ) r_\theta(x, y^*) r ΞΈ ​ ( x , y βˆ— ) is an unnormalized sum, its magnitude also scales linearly with N N N , quickly saturating the sigmoid. The derivative Οƒ β€² \sigma' Οƒ β€² drops toward zero, which would cause the optimizer to prematurely stop learning on long sequences. To convert this mechanism into the stable PCSS scaler, we make two fundamental mathematical transformations: Transformation 1: Sequence-Length Normalization PCSS explicitly normalizes the implied reward by sequence length using standard mean cross-entropy losses (where N = ∣ y βˆ— ∣ N = |y^*| N = ∣ y βˆ— ∣ ): Current model SFT loss : L sft ( ΞΈ ) = βˆ’ 1 N βˆ‘ t = 1 N log ⁑ Ο€ ΞΈ ( y t βˆ— ∣ x , y < t βˆ— ) L_{\text{sft}}(\theta) = -\frac{1}{N}\sum_{t=1}^{N} \log \pi_\theta(y_t^* \mid x, y_{<t}^*) L sft ​ ( ΞΈ ) = βˆ’ N 1 ​ βˆ‘ t = 1 N ​ lo g Ο€ ΞΈ ​ ( y t βˆ— ​ ∣ x , y < t βˆ— ​ ) Frozen base model reference loss : L ref = βˆ’ 1 N βˆ‘ t = 1 N log ⁑ Ο€ ref ( y t βˆ— ∣ x , y < t βˆ— ) L_{\text{ref}} = -\frac{1}{N}\sum_{t=1}^{N} \log \pi_{\text{ref}}(y_t^* \mid x, y_{<t}^*) L ref ​ = βˆ’ N 1 ​ βˆ‘ t = 1 N ​ lo g Ο€ ref ​ ( y t βˆ— ​ ∣ x , y < t βˆ— ​ ) The normalized reward becomes the difference: r norm = L ref βˆ’ L sft ( ΞΈ ) r_{\text{norm}} = L_{\text{ref}} - L_{\text{sft}}(\theta) r norm ​ = L ref ​ βˆ’ L sft ​ ( ΞΈ ) . Transformation 2: Dropping the KL Anchor ( z 0 = 0 ) (z_0 = 0) ( z 0 ​ = 0 ) The canonical KTO value function includes z 0 z_0 z 0 ​ as a baseline expectation ( z 0 = E y ∼ Ο€ ΞΈ ( β‹… ∣ x ) [ r ΞΈ ( x , y ) ] ) (z_0 = \mathbb{E}_{y \sim \pi_\theta(\cdot|x)}[r_\theta(x, y)]) ( z 0 ​ = E y ∼ Ο€ ΞΈ ​ ( β‹… ∣ x ) ​ [ r ΞΈ ​ ( x , y )]) . While KTO's batch-dependent z 0 z_0 z 0 ​ estimate breaks down at b s z = 1 bsz=1 b sz = 1 , mathematical analysis indicates that if we want the scaler to natively decay upon mastery in SFT, dropping the dynamic KL anchor is a straightforward architectural adjustment, making z 0 norm = 0 z_0^{\text{norm}} = 0 z 0 norm ​ = 0 a natural design choice . During fine-tuning on a ground-truth logic trace ( y βˆ— ) (y^*) ( y βˆ— ) , the policy's probability mass concentrates entirely on the correct answer ( Ο€ ΞΈ ( y βˆ— ∣ x ) β†’ 1 ) (\pi_\theta(y^*|x) \to 1) ( Ο€ ΞΈ ​ ( y βˆ— ∣ x ) β†’ 1 ) . If we plug this mastered state back into the formal sequence-normalized definition of z 0 z_0 z 0 ​ : z 0 norm = βˆ‘ y Ο€ ΞΈ ( y ∣ x ) [ 1 ∣ y ∣ log ⁑ Ο€ ΞΈ ( y ∣ x ) βˆ’ 1 ∣ y ∣ log ⁑ Ο€ ref ( y ∣ x ) ] β‰ˆ L ref z_0^{\text{norm}} = \sum_y \pi_\theta(y|x) \Big[ \frac{1}{|y|} \log \pi_\theta(y|x) - \frac{1}{|y|} \log \pi_{\text{ref}}(y|x) \Big] \approx L_{\text{ref}} z 0 norm ​ = y βˆ‘ ​ Ο€ ΞΈ ​ ( y ∣ x ) [ ∣ y ∣ 1 ​ lo g Ο€ ΞΈ ​ ( y ∣ x ) βˆ’ ∣ y ∣ 1 ​ lo g Ο€ ref ​ ( y ∣ x ) ] β‰ˆ L ref ​ Upon convergence ( L sft ( ΞΈ ) β†’ 0 ) (L_{\text{sft}}(\theta) \to 0) ( L sft ​ ( ΞΈ ) β†’ 0 ) , the reward is r norm = L ref βˆ’ L sft ( ΞΈ ) β‰ˆ L ref r_{\text{norm}} = L_{\text{ref}} - L_{\text{sft}}(\theta) \approx L_{\text{ref}} r norm ​ = L ref ​ βˆ’ L sft ​ ( ΞΈ ) β‰ˆ L ref ​ .
If we kept the baseline anchor ( z 0 norm ) (z_0^{\text{norm}}) ( z 0 norm ​ ) , the advantage would evaluate to: Advantage = r norm βˆ’ z 0 norm β‰ˆ L ref βˆ’ L ref = 0 \text{Advantage} = r_{\text{norm}} - z_0^{\text{norm}} \approx L_{\text{ref}} - L_{\text{ref}} = 0 Advantage = r norm ​ βˆ’ z 0 norm ​ β‰ˆ L ref ​ βˆ’ L ref ​ = 0 Because the derivative of the sigmoid ( Οƒ β€² ) (\sigma') ( Οƒ β€² ) reaches its peak value at an input of exactly 0 0 0 , keeping the baseline anchor ( z 0 norm ) (z_0^{\text{norm}}) ( z 0 norm ​ ) would cause the scaler to trend back toward its maximum value at the end of training , failing to suppress residual gradients and increasing the risk of late-stage overfitting. By explicitly setting z 0 norm = 0 z_0^{\text{norm}} = 0 z 0 norm ​ = 0 , the advantage becomes an absolute measure of progress against the base model ( L ref βˆ’ L sft ( ΞΈ ) ) (L_{\text{ref}} - L_{\text{sft}}(\theta)) ( L ref ​ βˆ’ L sft ​ ( ΞΈ )) . Upon mastery ( L sft ( ΞΈ ) β†’ 0 ) (L_{\text{sft}}(\theta) \to 0) ( L sft ​ ( ΞΈ ) β†’ 0 ) , the advantage becomes strictly positive ( L ref ) (L_{\text{ref}}) ( L ref ​ ) . The scaler structurally decays toward zero as mastery increases , naturally suppressing residual gradients to help mitigate late-stage overfitting. The Resulting Gradient The PCSS loss for a chosen example becomes: L chosen ( ΞΈ ) = 1 βˆ’ Οƒ ( Ξ² β‹… r norm ) \mathcal{L}_{\text{chosen}}(\theta) = 1 - \sigma(\beta \cdot r_{\text{norm}}) L chosen ​ ( ΞΈ ) = 1 βˆ’ Οƒ ( Ξ² β‹… r norm ​ ) βˆ‡ ΞΈ L chosen ( ΞΈ ) = [ Ξ² β‹… Οƒ β€² ( Ξ² β‹… r norm ) ] βˆ‡ ΞΈ L sft ( ΞΈ ) \nabla_\theta \mathcal{L}_{\text{chosen}}(\theta) = \Big[ \beta \cdot \sigma'(\beta \cdot r_{\text{norm}}) \Big] \nabla_\theta L_{\text{sft}}(\theta) βˆ‡ ΞΈ ​ L chosen ​ ( ΞΈ ) = [ Ξ² β‹… Οƒ β€² ( Ξ² β‹… r norm ​ ) ] βˆ‡ ΞΈ ​ L sft ​ ( ΞΈ ) Because this gradient mathematically reduces to the standard SFT gradient scaled by a dynamic multiplier, we can implement it directly by scaling the standard SFT loss output by the detached multiplier: L PCSS ( ΞΈ ) = L sft ( ΞΈ ) β‹… detach ( Ξ² β‹… Οƒ β€² ( Ξ² β‹… r norm ) ) \mathcal{L}_{\text{PCSS}}(\theta) = L_{\text{sft}}(\theta) \cdot \text{detach}(\beta \cdot \sigma'(\beta \cdot r_{\text{norm}})) L PCSS ​ ( ΞΈ ) = L sft ​ ( ΞΈ ) β‹… detach ( Ξ² β‹… Οƒ β€² ( Ξ² β‹… r norm ​ )) 2.2 Standard Deviation Normalization & Decoupled Scaling To standardize the dynamic multiplier across model families and decouple its magnitude from its shape, PCSS applies two scaling adjustments: Scaled Advantage : a = L ref βˆ’ L sft ( ΞΈ ) s global a = \frac{L_{\text{ref}} - L_{\text{sft}}(\theta)}{s_{\text{global}}} a = s global ​ L ref ​ βˆ’ L sft ​ ( ΞΈ ) ​ , where s global = Std D ( L ref ) s_{\text{global}} = \text{Std}_{\mathcal{D}}(L_{\text{ref}}) s global ​ = Std D ​ ( L ref ​ ) is the pre-computed standard deviation of reference losses across the training split. Decoupled Peak Scaling : Scale ( a ) = peak_scale β‹… ( 4.0 β‹… Οƒ β€² ( Ξ² β‹… a ) ) \text{Scale}(a) = \text{peak\_scale} \cdot (4.0 \cdot \sigma'(\beta \cdot a)) Scale ( a ) = peak_scale β‹… ( 4.0 β‹… Οƒ β€² ( Ξ² β‹… a )) Since max ⁑ ( 4.0 β‹… Οƒ β€² ) = 1.0 \max(4.0 \cdot \sigma') = 1.0 max ( 4.0 β‹… Οƒ β€² ) = 1.0 , peak_scale directly defines the maximum multiplier applied to the SFT gradient. Applying thes
TGSCrust367
🟧 echo.blog ⭐Reports mathematical reasoning gains after fine-tuning small base models on 100–500 zebra-puzzle traces without mathematical training exampltamewildβ€”β€”
🟧 hnCan My Puzzle Write-Ups 'Lift' Cheap LLMs?speckx10

Interpretation history

Decision trace