2026-10-11 17:20 UTC

The Heretic project claims its released tooling can remove refusal restrictions from supported open language models, potentially making unrestricted local variants easier to produce while weakening model-level safety controls.

state: resolvedheat: lowuncertainty: mediumknownscott: lowopen-models model-safety local-inferenceHeretic
Surfaced 2026-09-21T23:22:59Z β€” Heretic removes restrictions from language models β€” The episode now includes a separate released implementation claiming reversible KV-cache refusal suppression, extending the design space beyond Heretic's weight editing without validating Heretic itself. Cross-platform spread warrants high attention, but neither implementation has independent evidence here of reliable refusal removal with preserved capabilities.

What is this?

Heretic is an open-source CLI (pip: heretic-llm, AGPLv3) released in early 2026 by developer p-e-w (Philipp Emanuel Weidmann) that automatically removes refusal behavior β€” 'safety alignment' β€” from open-weight LLMs, building on Arditi et al. 2024's finding that refusal is mediated by a single residual-stream direction: directional ablation plus Optuna-driven parameter search that co-minimizes refusals and KL divergence from the base model. Widely cited benchmarks (Gemma-3-12B-IT: 3/100 refusals at KL 0.16 vs 0.45–1.04 for manual abliteration) all trace to the project's own eval tooling β€” presented as reproducible, but no snippet supplies genuinely independent audit, and coverage explicitly flags refusal-detection false positives as an open measurement problem. Traction is substantial: ~19.7K GitHub stars (+2.4K in the last week, top-15 trending), a reported 1,000+ community-published variants on Hugging Face, and the episode has widened into a family of alternative removal approaches (reversible KV-cache injection, activation steering) with mainstream amplification when PewDiePie covered the tool and shipped his own derived uncensored Qwen variant. Commentary also notes emerging hardening-against-ablation countermeasure research and the reverse application of ablating unwanted behaviors like sycophancy, while none of the supplied material settles the capability-preservation question beyond project-generated KL numbers.

Why it matters to Scott

Scott's own canon already carries the exact claim this case demonstrates β€” Architecture, Not Vibes and Guardrail Illusion hold that model-level behavioural controls are removable probability barriers and cannot serve as the enforcement boundary β€” so a pip-installable, mainstream-amplified removal tool is the world agreeing with him again, not news for him. The widened spread (PewDiePie as producer, daily-driver adoption of a Heretic finetune, a four-tool removal family) is material for the radar's proliferation tracking, but it challenges nothing in the hits and changes nothing he'd build or argue: removable vibes-level safety is the premise SiloOS and 'can't beats shouldn't' were designed around.
ip:framework.architecture-not-vibesip:concept.guardrail-illusionip:concept.trust-hierarchyradar:qwen38-abliteration-safety-tradeoffradar:gemma4-abliteration-evaluationradar:concept.activation-steering
queries asked of Scott's wikis
  • guardrail illusion model-level safety enforcement boundary
  • architecture not vibes behavioral controls removable
  • over-refusal blocking legitimate dev work local model workflow
  • open-weights local inference dual-use sovereignty regulation
  • activation steering refusal direction interpretability representation engineering
  • model eval harness capability preservation KL divergence refusal benchmarks

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

09-21 04:35⭐ origin directly observedHeretic removes restrictions from language models
Bluestein on hacker news
β€”
09-21 22:55first on r/LocalLLaMA Β· published Β· +18.3hUncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
Anony6666
β€”
09-24 14:33first on hacker news Β· published Β· +82.0hDynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
phatak-dev
β€”
10-02 17:25first on r/singularity Β· published Β· +276.8hPewDiePie says OpenAI banned his account(s) for distillation while building his AI model Ajax and experimenting with decrypting CoT
likeastar20
β€”
10-03 02:52first on r/artificial Β· published Β· +286.3hPewDiePie unveils "uncensored" Ajax AI model built to run on home PCs β€” creator says OpenAI banned him twice over model distillation used to build his product
ControlCAD
β€”
09-21 04:35amplified on hacker news πŸ‘‘hn.story.49783101
Bluestein
peak 279 Β· 111 comments Β· 25% of case engagement
09-21 22:55amplified on r/LocalLLaMAreddit.post.1wms904
Anony6666
peak 460 Β· 58 comments Β· 18% of case engagement
09-23 23:08amplified on r/LocalLLaMAreddit.post.1wol7zc
Khaledthe
peak 316 Β· 103 comments Β· 15% of case engagement
09-24 14:33amplified on hacker newshn.story.49831201
phatak-dev
peak 108 Β· 42 comments Β· 10% of case engagement
09-29 07:35amplified on hacker newshn.story.49889532
soltanov
peak 2 Β· 0 comments Β· 0% of case engagement
10-01 17:00amplified on r/LocalLLaMAreddit.post.1wv4vot
-p-e-w-
peak 314 Β· 89 comments Β· 14% of case engagement
3 more amplifiers in ainews.case_chain
09-21 10:20our radar first saw it Β· +5.8hdiscovery anchor: hn.story.49783101β€”
09-21 23:22reached heat=high Β· +18.8h Β· via ledgerβ€”β€”

Evidence (9) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Heretic removes restrictions from language models
Retrieved article excerpt

Open article Β· Retrieved 2026-09-21T10:22:32.208738+00:00

[Skip to content](https://heretic-project.org/#VPContent)

[Heretic](https://heretic-project.org/)

Search`⌘``Ctrl``K`

# Take control of the most important technology of our time

Heretic removes restrictions from language models, making sure they always follow your instructions

[GitHub](https://github.com/p-e-w/heretic)

[Hugging Face](https://huggingface.co/heretic-org)

[Discord](https://discord.gg/gdXc48gSyT)

[Matrix](https://matrix.to/#/#heretic:matrix.org)

Heretic logo

Get started in seconds:

sh

```
pip install -U heretic-llm
heretic Qwen/Qwen3.5-4B
```

[See the Tutorial >>](https://heretic-project.org/tutorial)
Bluestein279111
🟠 redditUncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
LocalLLaMA
Anony666646058
🟠 redditUsing uncensored models makes working less of a headache
LocalLLaMA
Khaledthe316103
🟧 hnDynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steeringphatak-dev10842
🟧 hnObliteratus Removes LLM Censorshipsoltanov20
🟠 redditHeretic is on PewDiePie!
LocalLLaMA
-p-e-w-31489
🟧 hnPewDiePie uncensors Qwen3.5-9B with Hereticalanwreath21
🟠 redditPewDiePie says OpenAI banned his account(s) for distillation while building his AI model Ajax and experimenting with decrypting CoT
singularity
likeastar208437
🟠 redditPewDiePie unveils "uncensored" Ajax AI model built to run on home PCs β€” creator says OpenAI banned him twice over model distillation used to build his product
artificial
ControlCAD42997

Interpretation history

Decision trace