Retrieved article excerpt
Open article Β· Retrieved 2026-09-11T14:22:15.938573+00:00
Back to Articles Eliciting Reasoning with 100-500 Zebra Puzzles Community Article Published
September 11, 2026 Upvote 1 tamewild tamewild Follow Resources & Quick Links: Read this Post: HuggingFace (Rendered Blog) | GitHub (Markdown Source) GitHub Repository: tamewild/pcss Reproduction Notebooks (Single H100/H200): Qwen 3 4B (~6.5 mins) | Granite 4.1 3B (~23 mins) | Qwen 3.5 9B (~40 mins) Models: PCSS-Qwen3-4B | PCSS-Granite-4.1-3B | PCSS-Qwen3.5-9B Datasets: zebra_100 | instruct5 1. Introduction Encountering the research on the Hyperfitting Phenomenon sparked an interest in exploring how pre-trained base models behave when micro-datasets are fine-tuned to near-zero training loss. We focused on training small pre-trained base models on 100 to 500 pure logic deduction traces (specifically 5x5 zebra puzzles) with zero mathematical data in the fine-tuning dataset . This synthetic dataset was generated by Qwen3-235B-A22B-Instruct-2507 in September 2025 using rejection sampling. Across multiple model families and architectures: Qwen 3 4B Base fine-tuned on 100 zebra puzzles (~6.5 minutes on 1 GPU) reaches 85.26% on the full 5,000-problem MATH benchmark (+31.16% delta over base) and 21.67% on AIME 2025 , rivaling the official post-trained model in non-thinking mode on MATH-500 while outperforming it on AIME 2025. Granite 4.1 3B Base fine-tuned on 500 zebra puzzles achieves 77.73% on MATH-500 and 19.44% on AIME 2025 , outperforming the official IBM Granite 4.1 3B Instruct model (+11.13 and +12.77 percentage points, respectively). Qwen 3.5 9B Base fine-tuned on 500 zebra puzzles jumps from 17.00% to 80.50% on 5x5 zebra puzzles (Pass@3: 97.05% ), while jumping from 5.42% to 22.92% on the uncontaminated ArXivMath 05/26 benchmark (Pass@3: 34.88% ) and reaching 96.60% on MATH-500 and 60.67% on AIME 2025 . While this out-of-domain generalization appears primarily driven by the pure logic dataset, we chose to manage the optimization process through adaptive loss scaling. For our primary results, we utilized PCSS (Per-Example Calibrated Sigmoid Scaler) : an adaptive, data-aware loss scaler derived directly from KTO. 2. The Mathematics of PCSS 2.1 From Preference Optimization to a Dynamic SFT Scaler PCSS originates from KTO . But why adapt an alignment algorithm for Supervised Fine-Tuning? The inspiration came from an experiment in the original KTO paper (Table 4). The authors demonstrated that running KTO using only desirable/positive examples (the exact same unpaired data used for SFT) yielded results highly competitive with standard SFT. While the paper conservatively notes this setup "yields similar results to SFT when Ξ² \beta Ξ² is small," the results with Ξ² = 0.1 \beta=0.1 Ξ² = 0.1 raised the overall average benchmark score from 29.7% to 31.1%, driven by a substantial relative increase on the GSM8K mathematical reasoning benchmark (from 1.0% to 12.5%). Why does an alignment algorithm behave this way? We can answer this by performing a gradient analysis on standard KTO when applied to purely desirable data. Setting the gain sensitivity to 1 ( Ξ» D = 1 ) (\lambda_D = 1) ( Ξ» D β = 1 ) , the KTO loss for a desirable ground-truth sequence y β y^* y β is: L KTO ( ΞΈ ) = 1 β Ο ( Ξ² ( r ΞΈ ( x , y β ) β z 0 ) ) \mathcal{L}_{\text{KTO}}(\theta) = 1 - \sigma(\beta(r_\theta(x, y^*) - z_0)) L KTO β ( ΞΈ ) = 1 β Ο ( Ξ² ( r ΞΈ β ( x , y β ) β z 0 β )) where r ΞΈ ( x , y ) r_\theta(x, y) r ΞΈ β ( x , y ) is the unnormalized log ratio ( log β‘ Ο ΞΈ ( y β£ x ) β log β‘ Ο ref ( y β£ x ) ) (\log \pi_\theta(y|x) - \log \pi_{\text{ref}}(y|x)) ( lo g Ο ΞΈ β ( y β£ x ) β lo g Ο ref β ( y β£ x )) . The loss for the ground-truth sequence is evaluated at r ΞΈ ( x , y β ) r_\theta(x, y^*) r ΞΈ β ( x , y β ) , and z 0 z_0 z 0 β is the KL divergence estimate. Crucially, KTO explicitly does not backpropagate through z 0 z_0 z 0 β . Taking the gradient with respect to ΞΈ \theta ΞΈ : β ΞΈ L KTO ( ΞΈ ) = β Ο β² ( Ξ² ( r ΞΈ ( x , y β ) β z 0 ) ) β
Ξ² β
β ΞΈ r ΞΈ ( x , y β ) \nabla_\theta \mathcal{L}_{\text{KTO}}(\theta) = -\sigma'(\beta(r_\theta(x, y^*) - z_0)) \cdot \beta \cdot \nabla_\theta r_\theta(x, y^*) β ΞΈ β L KTO β ( ΞΈ ) = β Ο β² ( Ξ² ( r ΞΈ β ( x , y β ) β z 0 β )) β
Ξ² β
β ΞΈ β r ΞΈ β ( x , y β ) Because the reference model is frozen, β ΞΈ r ΞΈ ( x , y β ) = β ΞΈ log β‘ Ο ΞΈ ( y β β£ x ) \nabla_\theta r_\theta(x, y^*) = \nabla_\theta \log \pi_\theta(y^*|x) β ΞΈ β r ΞΈ β ( x , y β ) = β ΞΈ β lo g Ο ΞΈ β ( y β β£ x ) . Since the standard SFT loss L sft ( ΞΈ ) L_{\text{sft}}(\theta) L sft β ( ΞΈ ) is the negative mean log-probability over N N N tokens, log β‘ Ο ΞΈ ( y β β£ x ) = β N β
L sft ( ΞΈ ) \log \pi_\theta(y^*|x) = -N \cdot L_{\text{sft}}(\theta) lo g Ο ΞΈ β ( y β β£ x ) = β N β
L sft β ( ΞΈ ) . Therefore, β ΞΈ log β‘ Ο ΞΈ ( y β β£ x ) = β N β
β ΞΈ L sft ( ΞΈ ) \nabla_\theta \log \pi_\theta(y^*|x) = -N \cdot \nabla_\theta L_{\text{sft}}(\theta) β ΞΈ β lo g Ο ΞΈ β ( y β β£ x ) = β N β
β ΞΈ β L sft β ( ΞΈ ) . Substituting this back in gives the exact gradient update for standard KTO on positive data: β ΞΈ L KTO ( ΞΈ ) = [ Ξ² β
Ο β² ( Ξ² ( r ΞΈ ( x , y β ) β z 0 ) ) β
N ] β ΞΈ L sft ( ΞΈ ) \nabla_\theta \mathcal{L}_{\text{KTO}}(\theta) = \Big[ \beta \cdot \sigma'(\beta(r_\theta(x, y^*) - z_0)) \cdot N \Big] \nabla_\theta L_{\text{sft}}(\theta) β ΞΈ β L KTO β ( ΞΈ ) = [ Ξ² β
Ο β² ( Ξ² ( r ΞΈ β ( x , y β ) β z 0 β )) β
N ] β ΞΈ β L sft β ( ΞΈ ) When you apply KTO to purely desirable examples, it is no longer performing preference alignment against negative data. Instead, it mathematically reduces to the standard SFT gradient multiplied by a dynamic scaler . However, this same equation highlights why standard KTO would struggle on logic chains spanning up to 11,000+ tokens: Length-Induced Gradient Inflation: The scaler is multiplied by the sequence length N N N . For long traces, the gradient magnitude increases sharply by a factor of hundreds. Length Bias & Gradient Saturation: Because r ΞΈ ( x , y β ) r_\theta(x, y^*) r ΞΈ β ( x , y β ) is an unnormalized sum, its magnitude also scales linearly with N N N , quickly saturating the sigmoid. The derivative Ο β² \sigma' Ο β² drops toward zero, which would cause the optimizer to prematurely stop learning on long sequences. To convert this mechanism into the stable PCSS scaler, we make two fundamental mathematical transformations: Transformation 1: Sequence-Length Normalization PCSS explicitly normalizes the implied reward by sequence length using standard mean cross-entropy losses (where N = β£ y β β£ N = |y^*| N = β£ y β β£ ): Current model SFT loss : L sft ( ΞΈ ) = β 1 N β t = 1 N log β‘ Ο ΞΈ ( y t β β£ x , y < t β ) L_{\text{sft}}(\theta) = -\frac{1}{N}\sum_{t=1}^{N} \log \pi_\theta(y_t^* \mid x, y_{<t}^*) L sft β ( ΞΈ ) = β N 1 β β t = 1 N β lo g Ο ΞΈ β ( y t β β β£ x , y < t β β ) Frozen base model reference loss : L ref = β 1 N β t = 1 N log β‘ Ο ref ( y t β β£ x , y < t β ) L_{\text{ref}} = -\frac{1}{N}\sum_{t=1}^{N} \log \pi_{\text{ref}}(y_t^* \mid x, y_{<t}^*) L ref β = β N 1 β β t = 1 N β lo g Ο ref β ( y t β β β£ x , y < t β β ) The normalized reward becomes the difference: r norm = L ref β L sft ( ΞΈ ) r_{\text{norm}} = L_{\text{ref}} - L_{\text{sft}}(\theta) r norm β = L ref β β L sft β ( ΞΈ ) . Transformation 2: Dropping the KL Anchor ( z 0 = 0 ) (z_0 = 0) ( z 0 β = 0 ) The canonical KTO value function includes z 0 z_0 z 0 β as a baseline expectation ( z 0 = E y βΌ Ο ΞΈ ( β
β£ x ) [ r ΞΈ ( x , y ) ] ) (z_0 = \mathbb{E}_{y \sim \pi_\theta(\cdot|x)}[r_\theta(x, y)]) ( z 0 β = E y βΌ Ο ΞΈ β ( β
β£ x ) β [ r ΞΈ β ( x , y )]) . While KTO's batch-dependent z 0 z_0 z 0 β estimate breaks down at b s z = 1 bsz=1 b sz = 1 , mathematical analysis indicates that if we want the scaler to natively decay upon mastery in SFT, dropping the dynamic KL anchor is a straightforward architectural adjustment, making z 0 norm = 0 z_0^{\text{norm}} = 0 z 0 norm β = 0 a natural design choice . During fine-tuning on a ground-truth logic trace ( y β ) (y^*) ( y β ) , the policy's probability mass concentrates entirely on the correct answer ( Ο ΞΈ ( y β β£ x ) β 1 ) (\pi_\theta(y^*|x) \to 1) ( Ο ΞΈ β ( y β β£ x ) β 1 ) . If we plug this mastered state back into the formal sequence-normalized definition of z 0 z_0 z 0 β : z 0 norm = β y Ο ΞΈ ( y β£ x ) [ 1 β£ y β£ log β‘ Ο ΞΈ ( y β£ x ) β 1 β£ y β£ log β‘ Ο ref ( y β£ x ) ] β L ref z_0^{\text{norm}} = \sum_y \pi_\theta(y|x) \Big[ \frac{1}{|y|} \log \pi_\theta(y|x) - \frac{1}{|y|} \log \pi_{\text{ref}}(y|x) \Big] \approx L_{\text{ref}} z 0 norm β = y β β Ο ΞΈ β ( y β£ x ) [ β£ y β£ 1 β lo g Ο ΞΈ β ( y β£ x ) β β£ y β£ 1 β lo g Ο ref β ( y β£ x ) ] β L ref β Upon convergence ( L sft ( ΞΈ ) β 0 ) (L_{\text{sft}}(\theta) \to 0) ( L sft β ( ΞΈ ) β 0 ) , the reward is r norm = L ref β L sft ( ΞΈ ) β L ref r_{\text{norm}} = L_{\text{ref}} - L_{\text{sft}}(\theta) \approx L_{\text{ref}} r norm β = L ref β β L sft β ( ΞΈ ) β L ref β .
If we kept the baseline anchor ( z 0 norm ) (z_0^{\text{norm}}) ( z 0 norm β ) , the advantage would evaluate to: Advantage = r norm β z 0 norm β L ref β L ref = 0 \text{Advantage} = r_{\text{norm}} - z_0^{\text{norm}} \approx L_{\text{ref}} - L_{\text{ref}} = 0 Advantage = r norm β β z 0 norm β β L ref β β L ref β = 0 Because the derivative of the sigmoid ( Ο β² ) (\sigma') ( Ο β² ) reaches its peak value at an input of exactly 0 0 0 , keeping the baseline anchor ( z 0 norm ) (z_0^{\text{norm}}) ( z 0 norm β ) would cause the scaler to trend back toward its maximum value at the end of training , failing to suppress residual gradients and increasing the risk of late-stage overfitting. By explicitly setting z 0 norm = 0 z_0^{\text{norm}} = 0 z 0 norm β = 0 , the advantage becomes an absolute measure of progress against the base model ( L ref β L sft ( ΞΈ ) ) (L_{\text{ref}} - L_{\text{sft}}(\theta)) ( L ref β β L sft β ( ΞΈ )) . Upon mastery ( L sft ( ΞΈ ) β 0 ) (L_{\text{sft}}(\theta) \to 0) ( L sft β ( ΞΈ ) β 0 ) , the advantage becomes strictly positive ( L ref ) (L_{\text{ref}}) ( L ref β ) . The scaler structurally decays toward zero as mastery increases , naturally suppressing residual gradients to help mitigate late-stage overfitting. The Resulting Gradient The PCSS loss for a chosen example becomes: L chosen ( ΞΈ ) = 1 β Ο ( Ξ² β
r norm ) \mathcal{L}_{\text{chosen}}(\theta) = 1 - \sigma(\beta \cdot r_{\text{norm}}) L chosen β ( ΞΈ ) = 1 β Ο ( Ξ² β
r norm β ) β ΞΈ L chosen ( ΞΈ ) = [ Ξ² β
Ο β² ( Ξ² β
r norm ) ] β ΞΈ L sft ( ΞΈ ) \nabla_\theta \mathcal{L}_{\text{chosen}}(\theta) = \Big[ \beta \cdot \sigma'(\beta \cdot r_{\text{norm}}) \Big] \nabla_\theta L_{\text{sft}}(\theta) β ΞΈ β L chosen β ( ΞΈ ) = [ Ξ² β
Ο β² ( Ξ² β
r norm β ) ] β ΞΈ β L sft β ( ΞΈ ) Because this gradient mathematically reduces to the standard SFT gradient scaled by a dynamic multiplier, we can implement it directly by scaling the standard SFT loss output by the detached multiplier: L PCSS ( ΞΈ ) = L sft ( ΞΈ ) β
detach ( Ξ² β
Ο β² ( Ξ² β
r norm ) ) \mathcal{L}_{\text{PCSS}}(\theta) = L_{\text{sft}}(\theta) \cdot \text{detach}(\beta \cdot \sigma'(\beta \cdot r_{\text{norm}})) L PCSS β ( ΞΈ ) = L sft β ( ΞΈ ) β
detach ( Ξ² β
Ο β² ( Ξ² β
r norm β )) 2.2 Standard Deviation Normalization & Decoupled Scaling To standardize the dynamic multiplier across model families and decouple its magnitude from its shape, PCSS applies two scaling adjustments: Scaled Advantage : a = L ref β L sft ( ΞΈ ) s global a = \frac{L_{\text{ref}} - L_{\text{sft}}(\theta)}{s_{\text{global}}} a = s global β L ref β β L sft β ( ΞΈ ) β , where s global = Std D ( L ref ) s_{\text{global}} = \text{Std}_{\mathcal{D}}(L_{\text{ref}}) s global β = Std D β ( L ref β ) is the pre-computed standard deviation of reference losses across the training split. Decoupled Peak Scaling : Scale ( a ) = peak_scale β
( 4.0 β
Ο β² ( Ξ² β
a ) ) \text{Scale}(a) = \text{peak\_scale} \cdot (4.0 \cdot \sigma'(\beta \cdot a)) Scale ( a ) = peak_scale β
( 4.0 β
Ο β² ( Ξ² β
a )) Since max β‘ ( 4.0 β
Ο β² ) = 1.0 \max(4.0 \cdot \sigma') = 1.0 max ( 4.0 β
Ο β² ) = 1.0 , peak_scale directly defines the maximum multiplier applied to the SFT gradient. Applying thes