2026-10-11 16:38 UTC

Prism ML (SkyIsNotGreen) claims its released Scion-35B-A3B โ€” a 35B-A3B MoE shipped as one 11.3GB GGUF with ternary PQ2_0 expert banks plus embedded trained corrections at 2.61 bpw, a bundled MTP drafter, and a required llama.cpp fork โ€” achieves Q4-class task retention at roughly half Q4_K_M's size, making ternary MoE a practically servable local tier; independent adoption and reproduced benchmarks confirm it, quiet fade closes it.

state: seedheat: lowuncertainty: highnovelscott: mediumternary-models local-inference mixture-of-experts quantizationPrism MLSkyIsNotGreen

What is this?

Per the case's anchor evidence (a Hugging Face repo titled 'SkyIsNotGreen/Scion-35B-A3B ยท Ternary MoE'), Prism ML โ€” publishing as SkyIsNotGreen โ€” has released Scion-35B-A3B, a 35B-total/~3B-active mixture-of-experts model shipped as a single 11.3GB GGUF with ternary PQ2_0 expert banks plus embedded trained corrections at 2.61 bpw, a bundled multi-token-prediction drafter, and a required llama.cpp fork, claiming Q4_K_M-class task quality at roughly half that format's size. The supplied search results do not surface this release at all โ€” no hit mentions Prism ML, SkyIsNotGreen, Scion, PQ2_0, ternary quantization, or the bpw/size claims โ€” so the release's specifics and any independent adoption are uncorroborated by this material. What the snippets do attest is a neighboring ecosystem data point: the LocalAI team's 'APEX' (Adaptive Precision for EXpert Models) quantizations of an unrelated 35B-A3B MoE bundle an MTP head into one GGUF for in-the-box self-speculative decoding via llama.cpp PR #22673 (--draft-mtp), with MoE-aware mixed-precision tiers and an MTP-aware imatrix patch still in progress. That mainstream pattern, however, ships Q3_K/Q4_K-class precision โ€” the snippets show the single-file MoE-plus-bundled-drafter delivery model is now standard tooling, but nothing in the supplied material corroborates the ternary 2-bit quality-retention claim at the heart of this case.

Why it matters to Scott

Novel first-party claim โ€” ternary PQ2_0 expert banks plus embedded trained corrections at 2.61 bpw hitting Q4-class retention in one 11.3GB GGUF โ€” is not a position Scott's wikis hold or argue against, so nothing converges or contradicts yet; but it lands directly on his stack: an 11.3GB/~3B-active resident MoE is exactly the always-on tier his gamepc/Ollama and OpenClaw setups would serve, while the required llama.cpp fork and custom quant type mean mainline Ollama/LM Studio can't load it today. Uncorroborated and quiet, so it changes nothing until the radar's adoption/reproduction condition fires โ€” validation would make it a live candidate for his local agent-model tier, fade would close it. Lineage to the Prism ML/Bonsai ternary line, the base3 packing sibling, and the bundled-MTP memory-overhead watch is recorded in radar_refs.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamadev:project.openclawradar:concept.ternary-modelsradar:concept.quantizationradar:bonsai-extreme-quantizationradar:bonsai-2-27b-ternary-releaseradar:base3-ternary-gguf-packingradar:llamacpp-fork-fragmentationradar:llama-cpp-mtp-default-memory-regression
queries asked of Scott's wikis
  • ternary 2-bit quantization viability for local serving
  • bpw quality retention claims vs Q4_K_M baseline methodology
  • llama.cpp fork requirement โ€” mainline vs fork tooling burden
  • bundled MTP drafter single-file GGUF speculative decoding
  • small active-parameter MoE for always-on local agent models
  • GGUF custom quant types and format fragmentation

Measured heat

now 0 pts/hpeak 11 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 143h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-05 16:41โญ origin directly observedSkyIsNotGreen/Scion-35B-A3B ยท Hugging Face - Ternary MoE
pmttyji on r/LocalLLaMA
โ€”
10-05 16:41amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wydetp
pmttyji
peak 18 ยท 11 comments ยท 100% of case engagement
10-05 17:20our radar first saw it ยท +0.7hdiscovery anchor: reddit.post.1wydetpโ€”
pace: p58 vs 1247 stories at the 96h mark (now 143h old) โ€” ahead of claude-auto-mode-classifier-outage (1.0x), behind ai-sre-arena-benchmark (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญSkyIsNotGreen/Scion-35B-A3B ยท Hugging Face - Ternary MoE
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-10-05T17:32:25.311259+00:00

# [SkyIsNotGreen](https://huggingface.co/SkyIsNotGreen) / [Scion-35B-A3B](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B) Like 9

[Text Generation](https://huggingface.co/models?pipeline_tag=text-generation)[GGUF](https://huggingface.co/models?library=gguf)[English](https://huggingface.co/models?language=en)[llama.cpp](https://huggingface.co/models?other=llama.cpp)[ternary](https://huggingface.co/models?other=ternary)[2-bit](https://huggingface.co/models?other=2-bit)[pq2\_0](https://huggingface.co/models?other=pq2_0)[llama-cpp](https://huggingface.co/models?other=llama-cpp)[Mixture of Experts](https://huggingface.co/models?other=moe)[qwen3.8](https://huggingface.co/models?other=qwen3.8)[quantized](https://huggingface.co/models?other=quantized)[scion](https://huggingface.co/models?other=scion)[speculative-decoding](https://huggingface.co/models?other=speculative-decoding)[draft-model](https://huggingface.co/models?other=draft-model)[mtp](https://huggingface.co/models?other=mtp)[conversational](https://huggingface.co/models?other=conversational)

License: apache-2.0

[Model card](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B)  [Files Files and versions  

xet](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/tree/main)  [Community 

3](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)

 

Deploy

  Copy to bucket new   

Use this model

 

[**GitHub: Build-scripts and docs**](https://github.com/sky-is-green/scion) ย |ย 
[**Forensics study**](https://github.com/sky-is-green/bonsai2-ternary-forensics) ย |ย 
[**Runtime fork**](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime) ย |ย 
[**Discussions**](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)

# Scion-35B-A3B: ternary MoE experts + trained corrections

A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus
small trained corrections, with no full-precision masters and no full-model QAT.

> **2.61 bpw** | **11.34 GB** (6.3x smaller than the BF16 reference) | **best PPL of the 2-bit class** | **task retention at Q4-class level, at about half of Q4\_K\_M's size**

The name comes from grafting. A scion is the shoot grafted onto a rootstock, and
here the trained corrections are grafted onto a 2.125 bpw ternary body.

## Highlights

- **One 11.34 GB file at 2.61 bpw** (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4\_K\_M is 21.71 GB and IQ2\_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
- **Task retention inside the Q4 and BF16 noise band**: HellaSwag 400 gives **79.00%** (BF16 81.25, Q4\_K\_M 80.00) and Winogrande gives **76.25%** (BF16 76.00, Q4\_K\_M 76.00). It has the joint-best Winogrande row in the table and the **best PPL of the 2-bit class** (8.354 against IQ2\_M 8.413 and Q2\_K 8.473).
- **Trained, not calibrated**: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
- **One file, no `--lora`**: the corrections are embedded (`adapter.embedded=true`) and attached at load. There is no adapter plumbing.
- **A k=1 speculative drafter ships alongside** ([`Scion-35B-A3B-mtp-drafter`](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter), 50 MB): it drafts the next token from the model's own hidden state and gives **1.14โ€“1.37ร— faster generation** on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).
- **The gap is stated, not hidden**: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4\_K\_M at 0.031. Tail-aware training was attempted at full scale and did not transfer ([`TAIL-EXPERIMENT-PLAN.md`](https://github.com/sky-is-green/scion/blob/main/docs/TAIL-EXPERIMENT-PLAN.md)).

## Resources

- **[GitHub `sky-is-green/scion`](https://github.com/sky-is-green/scion)**: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.
- **[Bonsai 2 ternary forensics](https://github.com/sky-is-green/bonsai2-ternary-forensics)**: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
- **Runtime**: [`sky-is-green/prism-ml-llama.cpp`](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime), branch `moe-corr-runtime`, a fork of Prism ML's llama.cpp with the PQ2\_0 container, the `ffn_moe_out` virtual target and embedded-adapter support.
- **[Retention grid](https://github.com/sky-is-green/scion/blob/main/retention-grid.png)**: this release against every community quant of the same model, same protocol.
- **[Discussions](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)**: questions, test reports and failures are all welcome.

## Model Overview

| Item | Specification |
| --- | --- |
| Base model | [`empero-ai/Qwen3.8-35B-A3B-Distill`](https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill), a Qwen3.8-line reasoning distill built on [`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0) |
| Parameters | 34.9B total, about 3B active per token (256 experts, top-8 plus shared) |
| Architecture | `qwen3_5_moe` (`qwen35moe` in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward |
| Context length | 262,144 tokens (inherited from the base model) |
| Weight format | Ternary **PQ2\_0** g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), **Q8\_0** for the rest, embedded corrections in the legacy `q1_0_g128` container (rank-512) |
| Low-bit coverage | Expert banks only; attention, embeddings and the head stay Q8\_0, and norms, routers and the output stay F32 |
| Deployed size | **10.558 GiB / 11.337 GB** (text only, single file) |
| Backends | llama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested |
| License | Apache-2.0 (inherited from the base) |

## Weight Representation: ternary PQ2\_0 + trained corrections

Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per
group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the
model is Q8\_0 (attention, embeddings, LM head) and F32 (norms, routers, output).
The corrections are rank-512 low-rank branches on the attention output and the
MoE block output, plus exact router deltas. They ship in the compact legacy
`q1_0_g128` container (2-bit codes plus one fp16 group scale per 128) and are
merged into the file. Effective overall: **2.61 bpw**.

### Memory Requirement

| Format | bpw | Size | vs BF16 |
| --- | --- | --- | --- |
| BF16 (reference) | 16.38 | 71.07 GB | 1.0x |
| Q8\_0 | 8.72 | 37.80 GB | 1.9x |
| Q4\_K\_M | 5.01 | 21.71 GB | 3.3x |
| IQ2\_M | 2.90 | 12.56 GB | 5.7x |
| **Scion-35B-A3B** | **2.61** | **11.34 GB** | **6.3x** |

Sizes are the published on-disk files of the same model, measured under one
protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see
Benchmarks).

### Shipped Components

| Component | Pack | Size | Residency |
| --- | --- | --- | --- |
| Language model (this repo) | PQ2\_0 experts, Q8\_0 rest, embedded corrections | 11.34 GB | resident; the whole model |
| **k=1 MTP drafter** ([separate repo](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter)) | **fp16 head (`fc1`/`gelu`/`fc2`) over the model's own hidden + next-token embedding** | **50 MB** | **transient; ~1 ms/eval on GPU** |
| Uncorrected body (not uploaded) | PQ2\_0 experts and Q8\_0 rest | 10.46 GiB | for swap tests |
| Corrections (not uploaded) | rank-512 branches and router deltas (`q1_0_g128`) | 98 MiB | for swap tests |

The language model is the single released body file; the drafter is a separate
50 MB file in its own repo
([`Scion-35B-A3B-mtp-drafter`](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter)).
The two-file variant (body plus separate adapter) exists for reproducing the
merge and swapping corrections at runtime. Ask in Discussions if you want it.

> **About the Hub's quant chip.** The Hub parses file names and labels this file
> `Q2_0`; the same happens on Prism ML's own `PQ2_0` releases. The container is
> Prism's **`PQ2_0`** (legacy name `Q1_0_g128`, type 142/43, identical byte
> layout): 2-bit codes with one fp16 group scale per 128 weights. It is not
> upstream llama.cpp's g64 `Q2_0` (type 42). No single quant name fits the file
> anyway, because it is a mix: `PQ2_0` expert banks, the embedded corrections in
> the legacy `q1_0_g128` container, and `Q8_0` for the rest (norms and routers
> in F32).

## Best Practices

### Generation Parameters

Recommended values, from the base model card:

> - `temperature=0.6`, `top_p=0.95`, `top_k=20`

This is a reasoning distill, and answers open with a long thinking segment.
Allow generous `max_new_tokens` (for example `-n 16384`); a small cap ends
generation mid-thought, before any answer.

### System Prompt

A simple prompt works, for example `You are a helpful assistant`. The base is a
reasoning SFT distill and does not require a special system prompt.

### Choosing Context and Offload

- The weights fit a 12 GB card; about 16 GB is comfortable with context.
- `-ngl 99` on a single card. When VRAM is tight, offload experts to CPU with
  `-ncmoe`, which is the VRAM-budget dial; the ternary container roughly halves
  the CPU-tail penalty compared to an f16 expert bank.
- **Threads should equal physical cores** (`-t 8` on an 8C/16T CPU; SMT siblings
  collapse CPU expert throughput).
- Do not layer-split across two cards when one card fits; the proxy
  measurements showed a 40% generation loss.

## Quickstart

> The runtime is the fork, and the fork is the source of truth for running these
> files: [`sky-is-green/prism-ml-llama.cpp`](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime),
> branch `moe-corr-runtime`.

### These files need the fork build

The PQ2\_0 container, the legacy `Q1_0_g128` import, the `ffn_moe_out` virtual
target and embedded adapters all live in the fork. **Stock llama.cpp will not
run this file**: it treats `PQ2_0` and `Q1_0_g128` as unknown tensor types.
Upstream's own `Q2_0` (type 42, g64) is a different container and is not a
substitute.
The fork's default branch is the one you want, so a plain clone is enough.

```
# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh          # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON     CPU-only: no flag
```

> **If the model will not load:** `tensor '...' has invalid ggml type 142. should be in [0, 43)` means the binary you ran is a build of upstream `master`, which
> knows 43 types and has never heard of `PQ2_0` (type id 142). The branch above
> defines `GGML_TYPE_PQ2_0 = 142` and `GGML_TYPE_COUNT = 144`, and
> `./verify-container-support.sh` checks exactly that in a second. The same error
> also comes from a stale `build/` directory or from an older `llama-cli` earlier
> on your `PATH`; the reliable fix for all three is to delete the checkout and
> the build directory and start from the clone above. See also the
> [fork README](https://github.com/sky-is-green/prism-ml-llama.cpp#readme).

```
# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
```

```
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -ngl 99 -c 4096 -t 8 \
    --temp 0.6 --top-p 0.95 --top-k 20 \
    -p "Explain quantum computing in simple
pmttyji1811

Interpretation history

Decision trace