2026-10-11 17:13 UTC

Independent evaluations will determine whether the abliterated Qwen3.8-27B FP8 checkpoint reduces harmful-request refusals to near zero while preserving general benchmark capability within roughly 1.3 points of the base model.

state: expiredheat: lowuncertainty: highknownscott: lowopen-models model-safety local-inferenceAlibabaQwen

What is this?

KridgeDookie has published a 27B-parameter multimodal derivative of Qwen/Qwen3.8-27B-FP8 that uses “abliteration” to suppress refusal behavior. Its deployment page claims zero refusals across two internal prompt sets, 23 of 24 coherence checks passed, and retained vision support; the supplied research snippet frames capability preservation as the key tradeoff in evaluating abliteration. The material does not provide an independent evaluation of the derivative or directly substantiate the claimed 0–6% refusal range and sub-1.3-point benchmark delta, so those remain model-card claims awaiting replication.

Why it matters to Scott

Scott’s Evaluation-Driven Development and independent-verification pages already require repeatable external evidence before accepting changed model behaviour; this case currently adds only unreplicated model-card claims. It is relevant to the radar’s established open-model evaluation and safety-capability tradeoff territory, but without independent results it would not yet change what Scott builds or argues.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.mechanically-different-verifiersradar:concept.open-modelsradar:concept.model-evaluationradar:concept.ai-safetyradar:compressed-llm-fidelity-safety-gap
queries asked of Scott's wikis
  • abliteration refusal vectors capability preservation
  • open-weight safety controls after weight release
  • uncensored local models safety asymmetry
  • independent evaluation of modified checkpoints
  • local inference model governance and sovereignty
  • benchmarking refusal reduction versus capability loss

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.8-27B abliterated FP8: refusal 64–99% → 0–6%, and MMLU/GSM8K move less than 1.3 points
LocalLLaMA
niacolhealth2312
🟧 echo.other ⭐The model card is the source of the Reddit post’s figures. It reports refusal falling from 64–99% for the official FP8 base to 0–6% after abOrcaRouter——
🟠 redditHuihui-ai Qwen 3.8 Ablit Available
LocalLLaMA
Frizzy-MacDrizzle2836
🟠 redditEccomerce frontend and a simple game one shooted by Qwen 3.8 27B
LocalLLaMA
Tiny-Assumption426397
🟠 redditLong Review: Qwen 3.8 27B is VERY good at tapping into it's real-world knowledge. It's "overthinking" brings it to Sonnet level performance with the potential for Opus level results.
LocalLLaMA
maxwell32142094

Interpretation history

Decision trace