KridgeDookie has published a 27B-parameter multimodal derivative of Qwen/Qwen3.8-27B-FP8 that uses “abliteration” to suppress refusal behavior. Its deployment page claims zero refusals across two internal prompt sets, 23 of 24 coherence checks passed, and retained vision support; the supplied research snippet frames capability preservation as the key tradeoff in evaluating abliteration. The material does not provide an independent evaluation of the derivative or directly substantiate the claimed 0–6% refusal range and sub-1.3-point benchmark delta, so those remain model-card claims awaiting replication.
Scott’s Evaluation-Driven Development and independent-verification pages already require repeatable external evidence before accepting changed model behaviour; this case currently adds only unreplicated model-card claims. It is relevant to the radar’s established open-model evaluation and safety-capability tradeoff territory, but without independent results it would not yet change what Scott builds or argues.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.mechanically-different-verifiersradar:concept.open-modelsradar:concept.model-evaluationradar:concept.ai-safetyradar:compressed-llm-fidelity-safety-gap
queries asked of Scott's wikis
- abliteration refusal vectors capability preservation
- open-weight safety controls after weight release
- uncensored local models safety asymmetry
- independent evaluation of modified checkpoints
- local inference model governance and sovereignty
- benchmarking refusal reduction versus capability loss
2026-08-17T15:35:25Z
Repeated comment refreshes have exhausted the near-term signal without producing any checkpoint-specific refusal test or controlled capability comparison. The case can fade from active tracking and be reopened if an independent evaluation appears.
2026-08-17T14:43:28Z
The comment refresh adds only more qualitative impressions of standard Qwen3.8 and no evidence tied to the abliterated FP8 checkpoint. Repeated amplification is not narrowing the uncertainty around refusal suppression or capability retention.
2026-08-17T14:08:55Z
The latest comment refresh is repetitive qualitative discussion of the standard Qwen3.8 model, not an independent test of the abliterated FP8 checkpoint. The claimed refusal reduction and capability retention remain uncorroborated.
2026-08-17T12:46:47Z
The refreshed comments are further qualitative amplification of standard Qwen3.8 capability, not checkpoint-specific evidence about the abliterated FP8 model. Without controlled refusal testing or comparative capability results, the claimed tradeoff remains uncorroborated.
2026-08-17T11:31:35Z
The refreshed comments remain qualitative impressions of the standard Qwen3.8 model and add no controlled test of the abliterated FP8 checkpoint. This is repetitive amplification, not evidence for the claimed refusal reduction or capability retention.
2026-08-17T10:35:13Z
The refreshed discussion remains qualitative commentary about the standard Qwen3.8 model and adds no checkpoint-specific refusal test, controlled capability comparison, or reproducible artifact. This is further repetitive amplification rather than independent corroboration of the claimed abliteration tradeoff.
2026-08-17T09:30:29Z
The refreshed comments remain conflicting, method-free impressions of the standard Qwen3.8 model rather than tests of the abliterated FP8 checkpoint. This is repetitive amplification, leaving the refusal-versus-capability claim uncorroborated.
2026-08-17T08:30:30Z
The new hands-on report strengthens the view that base Qwen3.8-27B is practically capable, but it evaluates a different quant and provides no checkpoint-specific refusal testing or controlled capability comparison. The abliterated model’s claimed safety-capability tradeoff therefore remains uncorroborated.
2026-08-17T08:22:28Z
evidence attached: reddit.post.1vqm51f — An independent hands-on report provides useful baseline evidence about Qwen 3.8 27B capability, though it does not directly test the abliterated checkpoint.
2026-08-17T05:27:06Z
The refreshed discussion remains qualitative commentary on base Qwen3.8 outputs and does not evaluate the abliterated FP8 checkpoint. With no checkpoint-specific refusal testing or controlled capability comparison, the claimed tradeoff remains uncorroborated.
2026-08-16T21:35:18Z
Refreshed comments merely critique and praise the base-model coding demo; they still do not identify the abliterated FP8 checkpoint or provide controlled refusal and capability measurements. The claimed safety-capability tradeoff remains uncorroborated.
2026-08-16T18:33:33Z
The coding demo shows that base Qwen3.8 can produce plausible local coding outputs, but it is not tied to the abliterated FP8 checkpoint and does not measure refusals or comparative benchmark retention. The central safety-capability tradeoff therefore remains an uncorroborated model-card claim.
2026-08-16T18:23:12Z
evidence attached: reddit.post.1vq37kt — The user's Qwen 3.8 coding examples provide limited qualitative context on practical local capability, though they do not test the safety tradeoff directly.
2026-08-16T17:40:08Z
Refreshed discussion adds source confusion and anecdotes about other HuiHui or Qwen variants, not an independent evaluation of this checkpoint. The refusal-versus-capability claim remains uncorroborated and dependent on the model card’s limited classifier.
2026-08-16T15:35:13Z
Refreshed comments add only private, method-free anecdotes that some abliterated variants perform well; they neither evaluate this checkpoint nor test its refusal metric. The case remains an uncorroborated model-card tradeoff claim awaiting reproducible independent results.
2026-08-16T13:26:40Z
The additional checkpoint-availability discussion confirms ecosystem interest but does not independently test the claimed refusal reduction or capability preservation. The case remains an uncorroborated model-card claim awaiting reproducible safety and benchmark evaluations.
2026-08-16T13:22:47Z
evidence attached: reddit.post.1vpw006 — This appears to corroborate the release of the Qwen 3.8 abliterated checkpoint, though it provides no independent capability or safety measurements.
2026-08-16T10:32:35Z
Refreshed discussion adds only anecdotal support that abliterated models can perform well on private coding tasks, without methods, results, or a reproducible artifact. The central refusal-versus-capability claim remains uncorroborated and dependent on the model card’s limited classifier.
2026-08-16T07:35:27Z
No independent evaluation or implementation evidence has arrived; the small engagement increase only amplifies the same self-reported, classifier-limited claims. The case remains a testable but uncorroborated safety-capability tradeoff.
2026-08-16T07:27:55Z
grounded: known/low — Scott’s Evaluation-Driven Development and independent-verification pages already require repeatable external evidence before accepting changed model behaviour;
2026-08-16T07:25:12Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vppox6 -> echo.other.e333fb1dc1 by OrcaRouter
2026-08-16T07:23:55Z
case created — The claimed refusal-versus-capability tradeoff is concrete and testable, but currently rests on one secondary report without a linked checkpoint or evaluation artifact.