Retrieved article excerpt
Open article ยท Retrieved 2026-10-05T17:32:25.311259+00:00
# [SkyIsNotGreen](https://huggingface.co/SkyIsNotGreen) / [Scion-35B-A3B](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B) Like 9
[Text Generation](https://huggingface.co/models?pipeline_tag=text-generation)[GGUF](https://huggingface.co/models?library=gguf)[English](https://huggingface.co/models?language=en)[llama.cpp](https://huggingface.co/models?other=llama.cpp)[ternary](https://huggingface.co/models?other=ternary)[2-bit](https://huggingface.co/models?other=2-bit)[pq2\_0](https://huggingface.co/models?other=pq2_0)[llama-cpp](https://huggingface.co/models?other=llama-cpp)[Mixture of Experts](https://huggingface.co/models?other=moe)[qwen3.8](https://huggingface.co/models?other=qwen3.8)[quantized](https://huggingface.co/models?other=quantized)[scion](https://huggingface.co/models?other=scion)[speculative-decoding](https://huggingface.co/models?other=speculative-decoding)[draft-model](https://huggingface.co/models?other=draft-model)[mtp](https://huggingface.co/models?other=mtp)[conversational](https://huggingface.co/models?other=conversational)
License: apache-2.0
[Model card](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B) [Files Files and versions
xet](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/tree/main) [Community
3](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)
Deploy
Copy to bucket new
Use this model
[**GitHub: Build-scripts and docs**](https://github.com/sky-is-green/scion) ย |ย
[**Forensics study**](https://github.com/sky-is-green/bonsai2-ternary-forensics) ย |ย
[**Runtime fork**](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime) ย |ย
[**Discussions**](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)
# Scion-35B-A3B: ternary MoE experts + trained corrections
A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus
small trained corrections, with no full-precision masters and no full-model QAT.
> **2.61 bpw** | **11.34 GB** (6.3x smaller than the BF16 reference) | **best PPL of the 2-bit class** | **task retention at Q4-class level, at about half of Q4\_K\_M's size**
The name comes from grafting. A scion is the shoot grafted onto a rootstock, and
here the trained corrections are grafted onto a 2.125 bpw ternary body.
## Highlights
- **One 11.34 GB file at 2.61 bpw** (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4\_K\_M is 21.71 GB and IQ2\_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
- **Task retention inside the Q4 and BF16 noise band**: HellaSwag 400 gives **79.00%** (BF16 81.25, Q4\_K\_M 80.00) and Winogrande gives **76.25%** (BF16 76.00, Q4\_K\_M 76.00). It has the joint-best Winogrande row in the table and the **best PPL of the 2-bit class** (8.354 against IQ2\_M 8.413 and Q2\_K 8.473).
- **Trained, not calibrated**: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
- **One file, no `--lora`**: the corrections are embedded (`adapter.embedded=true`) and attached at load. There is no adapter plumbing.
- **A k=1 speculative drafter ships alongside** ([`Scion-35B-A3B-mtp-drafter`](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter), 50 MB): it drafts the next token from the model's own hidden state and gives **1.14โ1.37ร faster generation** on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).
- **The gap is stated, not hidden**: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4\_K\_M at 0.031. Tail-aware training was attempted at full scale and did not transfer ([`TAIL-EXPERIMENT-PLAN.md`](https://github.com/sky-is-green/scion/blob/main/docs/TAIL-EXPERIMENT-PLAN.md)).
## Resources
- **[GitHub `sky-is-green/scion`](https://github.com/sky-is-green/scion)**: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.
- **[Bonsai 2 ternary forensics](https://github.com/sky-is-green/bonsai2-ternary-forensics)**: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
- **Runtime**: [`sky-is-green/prism-ml-llama.cpp`](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime), branch `moe-corr-runtime`, a fork of Prism ML's llama.cpp with the PQ2\_0 container, the `ffn_moe_out` virtual target and embedded-adapter support.
- **[Retention grid](https://github.com/sky-is-green/scion/blob/main/retention-grid.png)**: this release against every community quant of the same model, same protocol.
- **[Discussions](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B/discussions)**: questions, test reports and failures are all welcome.
## Model Overview
| Item | Specification |
| --- | --- |
| Base model | [`empero-ai/Qwen3.8-35B-A3B-Distill`](https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill), a Qwen3.8-line reasoning distill built on [`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0) |
| Parameters | 34.9B total, about 3B active per token (256 experts, top-8 plus shared) |
| Architecture | `qwen3_5_moe` (`qwen35moe` in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward |
| Context length | 262,144 tokens (inherited from the base model) |
| Weight format | Ternary **PQ2\_0** g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), **Q8\_0** for the rest, embedded corrections in the legacy `q1_0_g128` container (rank-512) |
| Low-bit coverage | Expert banks only; attention, embeddings and the head stay Q8\_0, and norms, routers and the output stay F32 |
| Deployed size | **10.558 GiB / 11.337 GB** (text only, single file) |
| Backends | llama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested |
| License | Apache-2.0 (inherited from the base) |
## Weight Representation: ternary PQ2\_0 + trained corrections
Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per
group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the
model is Q8\_0 (attention, embeddings, LM head) and F32 (norms, routers, output).
The corrections are rank-512 low-rank branches on the attention output and the
MoE block output, plus exact router deltas. They ship in the compact legacy
`q1_0_g128` container (2-bit codes plus one fp16 group scale per 128) and are
merged into the file. Effective overall: **2.61 bpw**.
### Memory Requirement
| Format | bpw | Size | vs BF16 |
| --- | --- | --- | --- |
| BF16 (reference) | 16.38 | 71.07 GB | 1.0x |
| Q8\_0 | 8.72 | 37.80 GB | 1.9x |
| Q4\_K\_M | 5.01 | 21.71 GB | 3.3x |
| IQ2\_M | 2.90 | 12.56 GB | 5.7x |
| **Scion-35B-A3B** | **2.61** | **11.34 GB** | **6.3x** |
Sizes are the published on-disk files of the same model, measured under one
protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see
Benchmarks).
### Shipped Components
| Component | Pack | Size | Residency |
| --- | --- | --- | --- |
| Language model (this repo) | PQ2\_0 experts, Q8\_0 rest, embedded corrections | 11.34 GB | resident; the whole model |
| **k=1 MTP drafter** ([separate repo](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter)) | **fp16 head (`fc1`/`gelu`/`fc2`) over the model's own hidden + next-token embedding** | **50 MB** | **transient; ~1 ms/eval on GPU** |
| Uncorrected body (not uploaded) | PQ2\_0 experts and Q8\_0 rest | 10.46 GiB | for swap tests |
| Corrections (not uploaded) | rank-512 branches and router deltas (`q1_0_g128`) | 98 MiB | for swap tests |
The language model is the single released body file; the drafter is a separate
50 MB file in its own repo
([`Scion-35B-A3B-mtp-drafter`](https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B-mtp-drafter)).
The two-file variant (body plus separate adapter) exists for reproducing the
merge and swapping corrections at runtime. Ask in Discussions if you want it.
> **About the Hub's quant chip.** The Hub parses file names and labels this file
> `Q2_0`; the same happens on Prism ML's own `PQ2_0` releases. The container is
> Prism's **`PQ2_0`** (legacy name `Q1_0_g128`, type 142/43, identical byte
> layout): 2-bit codes with one fp16 group scale per 128 weights. It is not
> upstream llama.cpp's g64 `Q2_0` (type 42). No single quant name fits the file
> anyway, because it is a mix: `PQ2_0` expert banks, the embedded corrections in
> the legacy `q1_0_g128` container, and `Q8_0` for the rest (norms and routers
> in F32).
## Best Practices
### Generation Parameters
Recommended values, from the base model card:
> - `temperature=0.6`, `top_p=0.95`, `top_k=20`
This is a reasoning distill, and answers open with a long thinking segment.
Allow generous `max_new_tokens` (for example `-n 16384`); a small cap ends
generation mid-thought, before any answer.
### System Prompt
A simple prompt works, for example `You are a helpful assistant`. The base is a
reasoning SFT distill and does not require a special system prompt.
### Choosing Context and Offload
- The weights fit a 12 GB card; about 16 GB is comfortable with context.
- `-ngl 99` on a single card. When VRAM is tight, offload experts to CPU with
`-ncmoe`, which is the VRAM-budget dial; the ternary container roughly halves
the CPU-tail penalty compared to an f16 expert bank.
- **Threads should equal physical cores** (`-t 8` on an 8C/16T CPU; SMT siblings
collapse CPU expert throughput).
- Do not layer-split across two cards when one card fits; the proxy
measurements showed a 40% generation loss.
## Quickstart
> The runtime is the fork, and the fork is the source of truth for running these
> files: [`sky-is-green/prism-ml-llama.cpp`](https://github.com/sky-is-green/prism-ml-llama.cpp/tree/moe-corr-runtime),
> branch `moe-corr-runtime`.
### These files need the fork build
The PQ2\_0 container, the legacy `Q1_0_g128` import, the `ffn_moe_out` virtual
target and embedded adapters all live in the fork. **Stock llama.cpp will not
run this file**: it treats `PQ2_0` and `Q1_0_g128` as unknown tensor types.
Upstream's own `Q2_0` (type 42, g64) is a different container and is not a
substitute.
The fork's default branch is the one you want, so a plain clone is enough.
```
# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON CPU-only: no flag
```
> **If the model will not load:** `tensor '...' has invalid ggml type 142. should be in [0, 43)` means the binary you ran is a build of upstream `master`, which
> knows 43 types and has never heard of `PQ2_0` (type id 142). The branch above
> defines `GGML_TYPE_PQ2_0 = 142` and `GGML_TYPE_COUNT = 144`, and
> `./verify-container-support.sh` checks exactly that in a second. The same error
> also comes from a stale `build/` directory or from an older `llama-cli` earlier
> on your `PATH`; the reliable fix for all three is to delete the checkout and
> the build directory and start from the clone above. See also the
> [fork README](https://github.com/sky-is-green/prism-ml-llama.cpp#readme).
```
# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
```
```
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-ngl 99 -c 4096 -t 8 \
--temp 0.6 --top-p 0.95 --top-k 20 \
-p "Explain quantum computing in simple