Retrieved article excerpt
Open article · Retrieved 2026-09-14T12:22:01.650615+00:00
### Uh oh!
There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).
[ggml-org](https://github.com/ggml-org)
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public
- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
23.2k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
128k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
# llama: add Maple 20B-A1B ternary MoE architecture (CPU) - #27000
#27000
Merged
[ggerganov](https://github.com/ggerganov) merged 8 commits into
[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom
[AlexGabbia:feat/maple-support](https://github.com/AlexGabbia/llama.cpp/tree/feat/maple-support)AlexGabbia/llama.cpp:feat/maple-supportCopy head branch name to clipboard
Sep 14, 2026
Merged
## [llama: add Maple 20B-A1B ternary MoE architecture (CPU)](https://github.com/ggml-org/llama.cpp/pull/27000#top)#27000 [ggerganov](https://github.com/ggerganov) merged 8 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [AlexGabbia:feat/maple-support](https://github.com/AlexGabbia/llama.cpp/tree/feat/maple-support)AlexGabbia/llama.cpp:feat/maple-supportCopy head branch name to clipboard
## Conversation
[@AlexGabbia](https://github.com/AlexGabbia)
### @AlexGabbia **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issue-5139870884) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).
Copy link
Copy Markdown
Contributor
## Overview
This adds the Maple 20B-A1B architecture to llama.cpp. Maple is DeepGrove's ternary MoE — 24 layers, 256 experts (8 active), SWA-512 interleaved with global attention at 3:1, ternary weights via TQ1\_0/TQ2\_0. Ported from `deepgrove-ai/llama.cpp` (commit `8ce8ca6c6d`) with their go-ahead (see [deepgrove-ai/llama.cpp#1](https://github.com/deepgrove-ai/llama.cpp/issues/1)).
Split into 4 commits: gguf-py constants → converter → architecture → test entry.
CPU-only for now — TQ kernels already exist in ggml. Metal and CUDA can follow.
## Additional information
- `test-llama-archs -a maple` passes (NMSE 8.75e-08)
- Smoke tested with `deepgrove/maple-preview` TQ2\_0 on M4 CPU: ~216 t/s prefill, ~88 t/s gen
- DeepGrove confirmed group size 256 works fine for this model (row-wise scales)
- Perplexity numbers being collected — will post results below
## Requirements
- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES — code was generated by an AI agent under my direction, based on DeepGrove's original implementation. I reviewed and tested everything before submitting.
[@AlexGabbia](https://github.com/AlexGabbia)
[AlexGabbia](https://github.com/AlexGabbia)
requested review from
[CISC](https://github.com/CISC),
[JohannesGaessler](https://github.com/JohannesGaessler) and
[ggerganov](https://github.com/ggerganov)
as [code owners](https://github.com/ggml-org/llama.cpp/blob/0d0bfcd4fd8828e3e7906b6fc4561725b534511e/CODEOWNERS#L29)
[August 13, 2026 09:28](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391183923)
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[model](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Amodel)
Model specific
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[conversion](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aconversion)
labels
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391370712)
[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
### This comment was marked as resolved.
[Sign in to view](https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fggml-org%2Fllama.cpp%2Fpull%2F27000)
[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
[ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
Bot
added
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391402748)
[@Green-Sky](https://github.com/Green-Sky)
### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5278779955) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).
Copy link
Copy Markdown
Collaborator
| |
| --- |
| I don't see `ternary TQ1_0/TQ2_0` being applicable, since they are group size 256, while the model was trained on 128. We still don't have 128 group side ternary, so only `q2_0` will work. ~~<https://huggingface.co/stamsam/maple-preview-gguf/blob/main/maple-q2_0.gguf>~~ |
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
marked this pull request as draft
[August 13, 2026 09:59](https://github.com/ggml-org/llama.cpp/pull/27000#event-29392589801)
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
removed
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29392591435)
[@AlexGabbia](https://github.com/AlexGabbia)
### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5279093578)
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| Thanks for the review. :) This PR supports the official DeepGrove GGUFs <https://huggingface.co/deepgrove/maple-preview-GGUF> which load and generate correctly on CPU — I tested the TQ2\_0 file on an Apple M4: 216 t/s prefill, 88 t/s generation. The stamsam GGUFs are built on a non-mainline fork (PrismML), so they don't help mainline support. I agree a 128-group ternary would be a good improvement and I'd be happy to work on it as a follow-up. |
[@Green-Sky](https://github.com/Green-Sky)
### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5280141305)
Copy link
Copy Markdown
Collaborator
| |
| --- |
| Please provide perplexity numbers, as suggested by the contributor guidelines. |
[@AlexGabbia](https://github.com/AlexGabbia)
### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5280500144)
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| Thanks — working on it. The official DeepGrove GGUFs include both TQ1\_0 and TQ2\_0 variants (with F16 and Q4\_K heads). Since Maple is natively ternary, the F16 GGUFs from DeepGrove are the same ternary model with F16 head/embeddings — not a full-precision baseline. I'll run perplexity comparing TQ2\_0 vs TQ1\_0 (both official DeepGrove GGUFs) and report the numbers here. |
[@Green-Sky](https://github.com/Green-Sky)
### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5281144950)
Copy link
Copy Markdown
Collaborator
| |
| --- |
| Someone with big internet please check if the conversion from source bf16 works, ideally [@AlexGabbia](https://github.com/AlexGabbia) . Some q8\_0 ggufs for comparison would be nice too, pretty sure q8\_0 or q4\_0 would be lossless (for the ternary tensors). |
[@Green-Sky](https://github.com/Green-Sky)
### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5282008849) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).
Copy link
Copy Markdown
Collaborator
| | | | | | | | | | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Ran a couple of ppl tests. I used their `TQ1_0-head-Q4_K` as base and quantized it to `q2_0` variants. `tq1_0` should convert to `q2_0` losslessly and `q2_0` runs on cuda. | quant | model size | ppl512 (default) | ppl2048 | | --- | --- | --- | --- | | `q2_0` `head-Q4_K` `embd-f16` | 6060.33 MiB (2.51 BPW) | 86.9194 +/- 0.89249 | 64.8277 +/- 0.68759 | | `q2_0` `head-q4_K` `embd-q8_0` | 5782.12 MiB (2.40 BPW) | 86.9098 +/- 0.89197 | 64.8001 +/- 0.68666 | | `q2_0` `head-q4_K` `embd-q6_K` | 5710.26 MiB (2.37 BPW) | 87.1487 +/- 0.89405 | 64.4658 +/- 0.68262 | The `q2_0` `head-Q4_K` `embd-f16` `gguf` should be lossless compared to base. Base size was 4747.45 MiB (1.97 BPW). Maple uses swa over 512 tokens in the sparse attending layers, so 2048 should capture that. Overall the ppl is suspiciously high, but that might be because of masked training, not sure. |
[@AlexGabbia](https://github.com/AlexGabbia)
### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5284791764) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).
Copy link
Copy Markdown
Contributor
Author
| | | | | | | | | | | | | | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Update with perplexity numbers — all 4 official DeepGrove GGUFs tested. **Setup**: wikitext-2, CPU M4 (10-core), 20 chunks, 8 threads, `-dev none` (CPU-only as this PR is CPU-scoped). | Model | Size | PPL [@512](https://github.com/512) | PPL [@2048](https://github.com/2048) | | --- | --- | --- | --- | | `tq2_0` head-F16 embd-f16 | 6,054 MB | 91.76 ± 5.15 | 56.65 ± 1.59 | | `tq2_0` head-Q4\_K embd-f16 | 5,902 MB | 94.15 ± 5.31 | 57.81 ± 1.63 | | `tq1_0` head-F16 embd-f16 | 5,431 MB | 91.76 ± 5.15 | 56.65 ± 1.59 | | `tq1_0` head-Q4\_K embd-f16 | 4,984 MB | 94.15 ± 5.31 | 57.81 ± 1.63 | **Conversion BF16 → GGUF F16**: works — 291 tensors, 40.5 GB, using the converter in this PR. **Observations:** - TQ1\_0 and TQ2\_0 produce identical PPL when using the same head type — expected, since both are packings of the same {-1,0,+1} ternary weights. - The [@2048](https://github.com/2048) PPL is significantly lower than [@512](https://github.com/512), consistent with Maple's SWA-512 architecture (as you noted). - The head-Q4\_K variant adds ~2.4 PPL points over head-F16, a small cost for ~150 MB size reduction. - All 4 official GGUFs use embd-f16 — the only difference between the F16 and Q4\_K variants is the output/head tensor type. - My numbers are in the same range as yours (86.92 [@512](https://github.com/512), 64.83 [@2048](https://github.com/2048)) — the difference is likely chunk count and sampling. **On the Q8\_0**: your q2\_0 head-Q4\_K embd-q8\_0 numbers already cover that comparison. The conversion from TQ1\_0 to q2\_0 is lossless for the ternary tensors as you said, and the embedding variants (F16/Q8\_0/Q6\_K) show negligible PPL difference — so I think that base is covered. |
[@Green-Sky](https://github.com/Green-Sky)
### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 14, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5291155467) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/g