2026-10-11 17:14 UTC

arbv claims the Unsloth-derived GPT-OSS chat template seriously degrades the model when replayed chat history contains prior analysis-channel reasoning and has published a corrected template โ€” upstream and derivative adoption would determine how widely default local GPT-OSS deployments are silently degraded.

state: seedheat: lowuncertainty: mediumconvergesscott: highgpt-oss chat-templates local-inference unslotharbvUnsloth

What is this?

gpt-oss is OpenAI's open-weight reasoning model family (20B/120B) built on the 'Harmony' multi-channel chat format (analysis/commentary/final channels), and most local deployments run through community-repackaged GGUFs whose Jinja chat templates are rewritten by distributors โ€” Unsloth documents finding multiple template bugs, fixing them in its own uploads (the default local distribution path per its docs), and pushing some fixes upstream to OpenAI's HF repos. arbv claims โ€” per the case's own evidence titles, which are the claim owner's announcement and corrected-template artifact โ€” that Unsloth's derived template degrades the model when replayed chat history contains prior analysis-channel reasoning, and has published a fixed template adding a preserve_thinking flag. The supplied snippets corroborate the surrounding frame: Unsloth actively patches gpt-oss templates, upstream HF discussions document recurring gpt-oss template errors (multi-turn handling, missing <|return|>, channel formatting), and a Gemma 4 discussion reports the same bug class of prior-turn reasoning re-injection, apparently via automated template audits โ€” but no snippet independently verifies arbv's specific degradation claim, which currently rests on his own artifact.

Why it matters to Scott

Converges with his provider-bound reasoning continuity concept and The Unverified Conversation ebook: preserve_thinking is precisely 'preserve reasoning replay within one compatible model loop,' and the alleged Unsloth template bug is the local client-asserted-history path failing that boundary while the model still appears functional โ€” silent drift in the default GGUF distribution chain. It bears on his own stack (Ollama on gamepc serving community GGUFs, gpt-oss being a likely local candidate he should template-check) and extends the radar's template-silently-breaks-reasoning family (Qwen3 thinking loss, Qwen3.8 default reasoning cost), with a dated-receipt opportunity for the ebook if the degradation reproduces beyond arbv's own artifact.
dev:concept.provider-bound-reasoning-continuityip:source.the-unverified-conversation-why-llms-can-t-trust-their-own-history-ebookip:concept.silent-driftdev:technology.ollamaradar:qwen3-chat-template-thinking-lossradar:qwen38-27b-default-reasoning-costradar:vllm-silent-tool-parser-failuresradar:ollama-silent-context-truncationradar:burrito-core-gpt-oss-harnessradar:concept.local-inferenceradar:concept.reasoning-tracesradar:concept.gguf
queries asked of Scott's wikis
  • chat template fidelity jinja llama.cpp vllm local inference
  • multi-turn replay of reasoning traces context assembly
  • harmony reasoning channel handling gpt-oss analysis commentary final
  • silent harness bugs mistaken for model quality degradation evals
  • unsloth gguf quant distribution adoption chain
  • agent memory persisting prior reasoning across turns

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 355h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-26 21:24 (minted)โญ origin echo-reconstructedUpdated GPT-OSS Jinja chat template adding preserve_thinking and fixing a bug induced by Unsloth's template that degrades the model when cha
arbv on github (echo) ยท attributed from reddit.post.1wr0wki ยท published time unknown
โ€”
09-26 20:35first on r/LocalLLaMA ยท published ยท lag ?Improved and fixed template for GPT-OSS (again). Includes preserve_thinking and fix for Unsloth-induced bug
arbv
โ€”
09-26 20:35amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wr0wki
arbv
peak 50 ยท 35 comments ยท 100% of case engagement
09-26 21:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wr0wkiโ€”
pace: p66 vs 1032 stories at the 336h mark (now 355h old) โ€” ahead of compute-cheap-h100-h200-pricing (1.0x), behind openai-german-wiki-incident (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditImproved and fixed template for GPT-OSS (again). Includes preserve_thinking and fix for Unsloth-induced bug
LocalLLaMA
arbv4735
๐ŸŸง echo.github โญUpdated GPT-OSS Jinja chat template adding preserve_thinking and fixing a bug induced by Unsloth's template that degrades the model when chaarbvโ€”โ€”

Interpretation history

Decision trace