2026-10-11 16:36 UTC

AlexGabbia claims the merged llama.cpp Maple 20B-A1B implementation runs DeepGrove's roughly 5–6 GB ternary GGUFs on CPU at about 88 generation tokens per second on an Apple M4, enabling GPU-free deployment while model quality remains uncertain.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumternary-quantization sparse-moe local-inferenceAlexGabbiaDeepGroveggml-orgGreen-Sky

What is this?

The supplied case identifies Maple 20B-A1B as DeepGrove’s ternary mixture-of-experts model and attributes a llama.cpp CPU-support implementation, PR #27000 in ggml-org/llama.cpp, to AlexGabbia. It reports roughly 5–6 GB GGUF files and Apple M4 CPU throughput of about 216 prefill and 88 generation tokens per second. None of the returned web snippets directly covers Maple or this pull request, so the claimed merge, model size and performance remain uncorroborated here; the snippets provide only background on CPU inference and other compact models. The supplied material does not establish Maple’s model quality or practical task performance.

Why it matters to Scott

The claimed CPU-only implementation converges with Scott’s Hardware-aware local inference approach and offers a concrete candidate to benchmark against his GPU-backed Ollama bulk classification/generation workloads, potentially changing their hardware requirements. That is a testing opportunity, not a demonstrated replacement: the merge and performance remain uncorroborated here, task quality is unknown, and the related radar cases do not establish prior tracking of Maple itself.
dev:concept.hardware-aware-local-inferencedev:technology.ollamaip:concept.ai-unit-economicsradar:concept.cpu-inferenceradar:concept.llama-cppradar:llama-cpp-bonsai-ternary-supportradar:moe-offload-constrained-cpu
queries asked of Scott's wikis
  • local inference economics GPU-free deployment
  • llama.cpp GGUF CPU inference projects
  • ternary quantization quality efficiency tradeoffs
  • sparse MoE active parameters memory footprint
  • local agent models task quality latency evaluation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1442h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-12 14:00⭐ origin echo-reconstructedAdds CPU support for DeepGrove's Maple ternary MoE, reports about 216 t/s prefill and 88 t/s generation on M4, and supplies conversion and p
AlexGabbia on github (echo) · attributed from reddit.post.1wg1o5b
—
09-14 12:11first on r/LocalLLaMA · published · +790.2hllama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
jacek2023
—
09-14 12:11amplified on r/LocalLLaMA 👑reddit.post.1wg1o5b
jacek2023
peak 156 · 46 comments · 100% of case engagement
09-14 12:20our radar first saw it · +790.3hdiscovery anchor: reddit.post.1wg1o5b—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditllama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-14T12:22:01.650615+00:00

### Uh oh!

There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).

[ggml-org](https://github.com/ggml-org) 
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public

- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
  23.2k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
   128k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)

# llama: add Maple 20B-A1B ternary MoE architecture (CPU) - #27000

#27000

Merged

[ggerganov](https://github.com/ggerganov) merged 8 commits into

[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom 

[AlexGabbia:feat/maple-support](https://github.com/AlexGabbia/llama.cpp/tree/feat/maple-support)AlexGabbia/llama.cpp:feat/maple-supportCopy head branch name to clipboard

Sep 14, 2026

Merged

## [llama: add Maple 20B-A1B ternary MoE architecture (CPU)](https://github.com/ggml-org/llama.cpp/pull/27000#top)#27000 [ggerganov](https://github.com/ggerganov) merged 8 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [AlexGabbia:feat/maple-support](https://github.com/AlexGabbia/llama.cpp/tree/feat/maple-support)AlexGabbia/llama.cpp:feat/maple-supportCopy head branch name to clipboard

## Conversation

[@AlexGabbia](https://github.com/AlexGabbia)

### @AlexGabbia **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issue-5139870884) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).

Copy link
 

Copy Markdown

Contributor

## Overview

This adds the Maple 20B-A1B architecture to llama.cpp. Maple is DeepGrove's ternary MoE — 24 layers, 256 experts (8 active), SWA-512 interleaved with global attention at 3:1, ternary weights via TQ1\_0/TQ2\_0. Ported from `deepgrove-ai/llama.cpp` (commit `8ce8ca6c6d`) with their go-ahead (see [deepgrove-ai/llama.cpp#1](https://github.com/deepgrove-ai/llama.cpp/issues/1)).

Split into 4 commits: gguf-py constants → converter → architecture → test entry.

CPU-only for now — TQ kernels already exist in ggml. Metal and CUDA can follow.

## Additional information

- `test-llama-archs -a maple` passes (NMSE 8.75e-08)
- Smoke tested with `deepgrove/maple-preview` TQ2\_0 on M4 CPU: ~216 t/s prefill, ~88 t/s gen
- DeepGrove confirmed group size 256 works fine for this model (row-wise scales)
- Perplexity numbers being collected — will post results below

## Requirements

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES — code was generated by an AI agent under my direction, based on DeepGrove's original implementation. I reviewed and tested everything before submitting.

[@AlexGabbia](https://github.com/AlexGabbia)

[AlexGabbia](https://github.com/AlexGabbia)
requested review from
[CISC](https://github.com/CISC),
[JohannesGaessler](https://github.com/JohannesGaessler) and
[ggerganov](https://github.com/ggerganov)
as [code owners](https://github.com/ggml-org/llama.cpp/blob/0d0bfcd4fd8828e3e7906b6fc4561725b534511e/CODEOWNERS#L29)
[August 13, 2026 09:28](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391183923)

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[model](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Amodel)
Model specific
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[conversion](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aconversion)
labels
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391370712)

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)

### This comment was marked as resolved.

[Sign in to view](https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fggml-org%2Fllama.cpp%2Fpull%2F27000)

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
[ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
Bot
added
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29391402748)

[@Green-Sky](https://github.com/Green-Sky)

### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5278779955) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).

Copy link
 

Copy Markdown

Collaborator

|  |
| --- |
| I don't see `ternary TQ1_0/TQ2_0` being applicable, since they are group size 256, while the model was trained on 128. We still don't have 128 group side ternary, so only `q2_0` will work.  ~~<https://huggingface.co/stamsam/maple-preview-gguf/blob/main/maple-q2_0.gguf>~~ |

[@github-actions](https://github.com/apps/github-actions)

[github-actions](https://github.com/apps/github-actions)
Bot
marked this pull request as draft
[August 13, 2026 09:59](https://github.com/ggml-org/llama.cpp/pull/27000#event-29392589801)

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
removed
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#event-29392591435)

[@AlexGabbia](https://github.com/AlexGabbia)

### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5279093578)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| Thanks for the review. :) This PR supports the official DeepGrove GGUFs <https://huggingface.co/deepgrove/maple-preview-GGUF> which load and generate correctly on CPU — I tested the TQ2\_0 file on an Apple M4: 216 t/s prefill, 88 t/s generation. The stamsam GGUFs are built on a non-mainline fork (PrismML), so they don't help mainline support. I agree a 128-group ternary would be a good improvement and I'd be happy to work on it as a follow-up. |

[@Green-Sky](https://github.com/Green-Sky)

### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5280141305)

Copy link
 

Copy Markdown

Collaborator

|  |
| --- |
| Please provide perplexity numbers, as suggested by the contributor guidelines. |

[@AlexGabbia](https://github.com/AlexGabbia)

### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5280500144)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| Thanks — working on it. The official DeepGrove GGUFs include both TQ1\_0 and TQ2\_0 variants (with F16 and Q4\_K heads). Since Maple is natively ternary, the F16 GGUFs from DeepGrove are the same ternary model with F16 head/embeddings — not a full-precision baseline. I'll run perplexity comparing TQ2\_0 vs TQ1\_0 (both official DeepGrove GGUFs) and report the numbers here. |

[@Green-Sky](https://github.com/Green-Sky)

### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5281144950)

Copy link
 

Copy Markdown

Collaborator

|  |
| --- |
| Someone with big internet please check if the conversion from source bf16 works, ideally [@AlexGabbia](https://github.com/AlexGabbia) .  Some q8\_0 ggufs for comparison would be nice too, pretty sure q8\_0 or q4\_0 would be lossless (for the ternary tensors). |

[@Green-Sky](https://github.com/Green-Sky)

### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5282008849) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).

Copy link
 

Copy Markdown

Collaborator

|  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Ran a couple of ppl tests.  I used their `TQ1_0-head-Q4_K` as base and quantized it to `q2_0` variants. `tq1_0` should convert to `q2_0` losslessly and `q2_0` runs on cuda.   | quant | model size | ppl512 (default) | ppl2048 | | --- | --- | --- | --- | | `q2_0` `head-Q4_K` `embd-f16` | 6060.33 MiB (2.51 BPW) | 86.9194 +/- 0.89249 | 64.8277 +/- 0.68759 | | `q2_0` `head-q4_K` `embd-q8_0` | 5782.12 MiB (2.40 BPW) | 86.9098 +/- 0.89197 | 64.8001 +/- 0.68666 | | `q2_0` `head-q4_K` `embd-q6_K` | 5710.26 MiB (2.37 BPW) | 87.1487 +/- 0.89405 | 64.4658 +/- 0.68262 |   The `q2_0` `head-Q4_K` `embd-f16` `gguf` should be lossless compared to base.  Base size was 4747.45 MiB (1.97 BPW).  Maple uses swa over 512 tokens in the sparse attending layers, so 2048 should capture that.  Overall the ppl is suspiciously high, but that might be because of masked training, not sure. |

[@AlexGabbia](https://github.com/AlexGabbia)

### **[AlexGabbia](https://github.com/AlexGabbia)** commented [Aug 13, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5284791764) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27000).

Copy link
 

Copy Markdown

Contributor

Author

|  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Update with perplexity numbers — all 4 official DeepGrove GGUFs tested.  **Setup**: wikitext-2, CPU M4 (10-core), 20 chunks, 8 threads, `-dev none` (CPU-only as this PR is CPU-scoped).   | Model | Size | PPL [@512](https://github.com/512) | PPL [@2048](https://github.com/2048) | | --- | --- | --- | --- | | `tq2_0` head-F16 embd-f16 | 6,054 MB | 91.76 ± 5.15 | 56.65 ± 1.59 | | `tq2_0` head-Q4\_K embd-f16 | 5,902 MB | 94.15 ± 5.31 | 57.81 ± 1.63 | | `tq1_0` head-F16 embd-f16 | 5,431 MB | 91.76 ± 5.15 | 56.65 ± 1.59 | | `tq1_0` head-Q4\_K embd-f16 | 4,984 MB | 94.15 ± 5.31 | 57.81 ± 1.63 |   **Conversion BF16 → GGUF F16**: works — 291 tensors, 40.5 GB, using the converter in this PR.  **Observations:**   - TQ1\_0 and TQ2\_0 produce identical PPL when using the same head type — expected, since both are packings of the same {-1,0,+1} ternary weights. - The [@2048](https://github.com/2048) PPL is significantly lower than [@512](https://github.com/512), consistent with Maple's SWA-512 architecture (as you noted). - The head-Q4\_K variant adds ~2.4 PPL points over head-F16, a small cost for ~150 MB size reduction. - All 4 official GGUFs use embd-f16 — the only difference between the F16 and Q4\_K variants is the output/head tensor type. - My numbers are in the same range as yours (86.92 [@512](https://github.com/512), 64.83 [@2048](https://github.com/2048)) — the difference is likely chunk count and sampling.   **On the Q8\_0**: your q2\_0 head-Q4\_K embd-q8\_0 numbers already cover that comparison. The conversion from TQ1\_0 to q2\_0 is lossless for the ternary tensors as you said, and the embedding variants (F16/Q8\_0/Q6\_K) show negligible PPL difference — so I think that base is covered. |

[@Green-Sky](https://github.com/Green-Sky)

### **[Green-Sky](https://github.com/Green-Sky)** commented [Aug 14, 2026](https://github.com/ggml-org/llama.cpp/pull/27000#issuecomment-5291155467) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/g
jacek202315646
🟧 echo.github ⭐Adds CPU support for DeepGrove's Maple ternary MoE, reports about 216 t/s prefill and 88 t/s generation on M4, and supplies conversion and pAlexGabbia——

Interpretation history

Decision trace