2026-10-11 18:03 UTC

Independent review will determine whether the reported Gemma 4 12B abliteration methods materially reduce refusals without unacceptable reasoning, benchmark, or output-quality losses.

state: expiredheat: lowuncertainty: highknownscott: lowmodel-editing open-models ai-safety-evaluationGoogleNathan Dreamfast

What is this?

Gemma 4 12B is described by Google DeepMind as an open-weight, encoder-free multimodal model designed for local hardware. GitHub and Hugging Face pages report abliteration methods that reduce its refusal rate to near zero while retaining benchmark parity or low KL divergence, but these are project or model-host claims rather than clearly independent evaluations. The supplied independent-looking benchmark concerns Gemma 3, not Gemma 4, and the snippets do not establish Nathan Dreamfast’s role or independently verify the reported 165 GPU-hour experiment and absence of broader reasoning or output-quality losses.

Why it matters to Scott

Scott’s “Evaluation-Driven Development” already requires repeatable quality gates before shipping changed model behaviour, including checks for capability and output-quality regressions; the radar also tracks the nearly identical abliteration trade-off in “Independent evaluations will determine whether the abliterated Qwen3.8-27B…” This Gemma checkpoint is another unverified instance of that established question, with no supplied independent result or evidence that it affects a model Scott actively uses.
ip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersradar:qwen38-abliteration-safety-tradeoffradar:concept.model-safetyradar:concept.llm-evaluationradar:concept.open-models
queries asked of Scott's wikis
  • abliteration and refusal-direction model editing
  • capability preservation after safety-weight modification
  • open-weight models and downstream alignment control
  • refusal benchmarks versus harmful-capability evaluations
  • local model sovereignty and removable safeguards
  • evaluation design for reasoning and output-quality regressions

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐12 abliterated Gemma 4 12B variants, one base, 165 GPU hours - Abliterlitics
LocalLLaMA
nathandreamfast5318

Interpretation history

Decision trace