Heretic is an open-source CLI (pip: heretic-llm, AGPLv3) released in early 2026 by developer p-e-w (Philipp Emanuel Weidmann) that automatically removes refusal behavior β 'safety alignment' β from open-weight LLMs, building on Arditi et al. 2024's finding that refusal is mediated by a single residual-stream direction: directional ablation plus Optuna-driven parameter search that co-minimizes refusals and KL divergence from the base model. Widely cited benchmarks (Gemma-3-12B-IT: 3/100 refusals at KL 0.16 vs 0.45β1.04 for manual abliteration) all trace to the project's own eval tooling β presented as reproducible, but no snippet supplies genuinely independent audit, and coverage explicitly flags refusal-detection false positives as an open measurement problem. Traction is substantial: ~19.7K GitHub stars (+2.4K in the last week, top-15 trending), a reported 1,000+ community-published variants on Hugging Face, and the episode has widened into a family of alternative removal approaches (reversible KV-cache injection, activation steering) with mainstream amplification when PewDiePie covered the tool and shipped his own derived uncensored Qwen variant. Commentary also notes emerging hardening-against-ablation countermeasure research and the reverse application of ablating unwanted behaviors like sycophancy, while none of the supplied material settles the capability-preservation question beyond project-generated KL numbers.
Scott's own canon already carries the exact claim this case demonstrates β Architecture, Not Vibes and Guardrail Illusion hold that model-level behavioural controls are removable probability barriers and cannot serve as the enforcement boundary β so a pip-installable, mainstream-amplified removal tool is the world agreeing with him again, not news for him. The widened spread (PewDiePie as producer, daily-driver adoption of a Heretic finetune, a four-tool removal family) is material for the radar's proliferation tracking, but it challenges nothing in the hits and changes nothing he'd build or argue: removable vibes-level safety is the premise SiloOS and 'can't beats shouldn't' were designed around.
ip:framework.architecture-not-vibesip:concept.guardrail-illusionip:concept.trust-hierarchyradar:qwen38-abliteration-safety-tradeoffradar:gemma4-abliteration-evaluationradar:concept.activation-steering
queries asked of Scott's wikis
- guardrail illusion model-level safety enforcement boundary
- architecture not vibes behavioral controls removable
- over-refusal blocking legitimate dev work local model workflow
- open-weights local inference dual-use sovereignty regulation
- activation steering refusal direction interpretability representation engineering
- model eval harness capability preservation KL divergence refusal benchmarks
2026-10-04T07:52:13Z
Episode closed as established: every velocity_spike trigger since the crest was a stale residual β total-score crossings of the p90 line by 300h-old posts (1wv4vot at 314, the 1wwccsz Ajax repost at 244β310) plus a +4-comment drift β with aggregate velocity at ~4.3 pts/h against a 305 peak, zero new evidence, and the periphery no longer expanding (no new tools, communities or outlets since the PewDiePie wave crested; the 87.5 peer percentile is one already-tracked repost accruing sunk upvotes). The removal-tool ecosystem is now landscape fact rather than a live story; the predicted adoption flood and regulatory response never fired, and future signal (independent audit, hardening countermeasures, regulation) belongs to the sibling radar entries or a fresh case.
2026-10-03T05:37:10Z
The PewDiePie wave has crested: the newest attachment is a near-zero-engagement repost of the already-tracked Ajax/distillation story, not a new community, implementation or outlet, and the velocity_spike trigger reads stale β residual upvotes on a 288h-old post, with aggregate velocity halved again to ~1.8 pts/h against a 277 peak. The case's meaning is unchanged (an established removal-tool ecosystem resting on still-unverified effectiveness numbers), so it cools to low heat: the magnitude-valve reading reflects sunk launch-coverage totals, not ongoing spread, since the periphery has stopped expanding and the adoption flood remains unfired.
2026-10-03T03:23:50Z
evidence attached: reddit.post.1wwccsz β shared external link with case evidence
2026-10-02T18:50:22Z
The newly attached PewDiePie OpenAI-distillation-ban thread is adjacent fallout, not Heretic evidence: it explains why a mainstream creator pivoted from frontier distillation to local open weights and thence to Heretic, and reaches this case only through the shared actor. The case's meaning is unchanged β a widening removal-tool periphery (four tools, daily-driver adoption, PewDiePie now a producer) around an entirely unverified effectiveness claim β so state, uncertainty and low Scott-relevance all hold. Medium heat is kept despite a cooling ~4.3 pts/h aggregate because the periphery is still emitting top-decile objects (this post, 91.7 cohort percentile) within the first day of the PewDiePie wave; all escalation triggers remain unfired.
2026-10-02T17:39:06Z
evidence attached: reddit.post.1wvzpl7 β shared external link with case evidence
2026-10-02T16:47:56Z
grounded: known/low β Scott's own canon already carries the exact claim this case demonstrates β Architecture, Not Vibes and Guardrail Illusion hold that model-level behavioural cont
2026-10-02T16:39:37Z
PewDiePie graduates from amplifier to producer: a derived uncensored Qwen3.5-9B built with Heretic is claimed via HN (2 pts, 1 comment β unverified, but he is a rooted local-LLM creator), firing the case's 'influential entrant ships a derived model' trigger in miniature. State advances to accelerating on the widening periphery (four removal tools, daily-driver adoption, a mainstream creator now shipping variants), while heat holds at medium because the predicted adoption flood has not landed and measured velocity is a cooling ~5 pts/h tail β one Reddit meta-thread re-spiking at ~14 pts/h against a near-zero 276h baseline, not platform dominance.
2026-10-02T16:26:23Z
evidence attached: hn.story.49934603 β High-profile creator shipping an uncensored Qwen variant via Heretic is adoption-and-spread evidence for the unrestriction-tooling case.
2026-10-01T19:07:50Z
PewDiePie coverage moves Heretic from practitioner-niche tooling into mainstream awareness: the case now tracks a removal tool with mass-distribution potential, making the adoption-flood and regulatory-backlash scenarios materially more live while leaving effectiveness and capability-damage claims entirely unverified. Held at medium rather than high because current velocity (~9 pts/h vs a 207 peak) is a decayed tail plus one modest meta-thread, and the magnitude-valve spread reading was already priced at the September peak; escalation waits for the predicted flood of new users and derived models to actually land.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv4vot β PewDiePie coverage is mainstream-scale amplification of the restriction-removal tool, materially changing the spread the open case tracks.
2026-09-29T09:00:59Z
Obliteratus arrives as a fourth independent refusal-removal tool, but only as a title-only, zero-traction HN submission β nothing in the item supports the 'prominent actor' framing in the attach rationale β so it firms the spreading-family reading without adding any effectiveness or adoption evidence. With engagement fully decayed (165β~0.3 pts/h, phantom-kv's cohort percentile down from 92nd to 76th) and the only new periphery a noise-level entrant, this is a cooling of attention, not belief: the hypothesis stays corroborated and still unverified in effect.
2026-09-29T08:25:12Z
evidence attached: hn.story.49889532 β A second refusal-removal tool from a different prominent actor extends the open case from one project to a spreading pattern.
2026-09-24T17:49:10Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-24T16:38:17Z
evidence attached: hn.story.49831201 β Independent non-destructive steering-based refusal suppression corroborates the case's claim that model-level safety controls are increasingly easy to strip.
2026-09-24T00:41:48Z
An independent user now reports daily-driving a Heretic-derived model for ordinary work, providing initial adoption evidence beyond the projectβs own release claims. This corroborates proliferation of easily accessible refusal-suppressed variants, but not reliable safety removal or capability preservation.
2026-09-24T00:31:34Z
evidence attached: reddit.post.1wol7zc β Real adoption evidence: a Heretic-produced unrestricted finetune used as a daily work driver because base models refuse ordinary tasks, supporting the case's proliferation/safety-erosion hypothesis.
2026-09-21T23:22:59Z
The episode now includes a separate released implementation claiming reversible KV-cache refusal suppression, extending the design space beyond Heretic's weight editing without validating Heretic itself. Cross-platform spread warrants high attention, but neither implementation has independent evidence here of reliable refusal removal with preserved capabilities.
2026-09-21T23:22:20Z
evidence attached: reddit.post.1wms904 β The released phantom-kv artifact extends the same model-unrestriction episode with a reversible, hot-swappable KV-cache approach.
2026-09-21T13:58:58Z
grounded: known/low β Scott already argues that model-level behavioural safety is removable and cannot serve as the enforcement boundary in Architecture, Not Vibes and Guardrail Illu
2026-09-21T13:55:48Z
case created β The installable artifact has immediate operational and safety implications and has attracted material early attention.