ggml-org (Xuan-Son Nguyen, Victor Mustar) shipped a native `/v1/systemone` endpoint in llama.cpp server (PR #29818, announced Oct 2 2026) that serves an open collection of five decision models — Julia-1 144M, Laya 421M, Kev-4B, lev 4B, and vision-capable OpenJev 27B — returning per-option probabilities in a single forward pass with zero generated tokens. The System One wire format (state + typed questions: choice/score/noul) originated from TypeSafe's Jev; llama.cpp's implementation uses existing GBNF grammar, n_probs, and prompt-cache primitives. Multi-runtime adoption followed within days: Ollama native `/v1/systemone` (v0.35, Sept 29), vLLM-side 'Decision 2.0' models (Oct 4-5), and AWS-backed Strands shipping Decider 2B as the first mainstream cloud-vendor agent-stack entrant. Independent ecosystem signals include TinyDecide 10M (browser→ESP32), Jiwo fine-tunes topping the Decision Index, a shipped browser agent using Jev, Pi coding agent plugin using Jev/Kev/Laya for command classification, and multiple independent benchmarks via the llama.cpp endpoint. The weekly open-model cadence commitment is being honored: Cloudflare Clef verified working on the endpoint via bartowski quant (Oct 8).
| source | object | author | score | comments |
| 🟠 reddit | New in llama.cpp: Decision Models LocalLLaMA Retrieved article excerptOpen article · Retrieved 2026-10-02T15:25:30.921317+00:00 [Back to Articles](https://huggingface.co/blog)
# New in llama.cpp: Decision Models
[Community Article](https://huggingface.co/blog/community)
Published
October 2, 2026
[Upvote
14](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp)
- +8
[Xuan-Son Nguyen's avatar](https://huggingface.co/ngxson)
[Xuan-Son Nguyen
ngxson
Follow](https://huggingface.co/ngxson)
[ggml-org's avatar](https://huggingface.co/ggml-org "ggml-org") [ggml-org](https://huggingface.co/ggml-org)
[Victor Mustar's avatar](https://huggingface.co/victor)
[Victor Mustar
victor
Follow](https://huggingface.co/victor)
[ggml-org's avatar](https://huggingface.co/ggml-org "ggml-org") [ggml-org](https://huggingface.co/ggml-org)
llama.cpp server now supports decision models through the `/v1/systemone` endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.
The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in [PR #29818](https://github.com/ggml-org/llama.cpp/pull/29818).
> **What is a decision model?** A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.
## Supported models
| Model | Size | Based on | Languages | Images | License | Speed\* |
| --- | --- | --- | --- | --- | --- | --- |
| [Julia-1](https://huggingface.co/ggml-org/Julia-1-GGUF) | 144M | mmBERT-small | 50+ | no | Apache 2.0 | 3 ms |
| [Laya](https://huggingface.co/ggml-org/Laya-GGUF) | 421M | ModernBERT-large | English | no | Apache 2.0 | 5 ms |
| [Kev-4B](https://huggingface.co/ggml-org/Kev-4B-GGUF) | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | 12 ms |
| [lev](https://huggingface.co/ggml-org/lev-GGUF) | 4B | Qwen3.5-4B | English | no | Apache 2.0 | 36 ms |
| [OpenJev](https://huggingface.co/ggml-org/OpenJev-GGUF) | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | 43 ms |
\*Median time to answer one question, on one NVIDIA RTX PRO 6000.
Find these models in the [Decision models collection](https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769), with more coming. The community [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) shows how they compare.
## Quick start
Get the latest llama.cpp from [llama.app](https://llama.app) (or run `llama update`), then start a model:
```
llama serve -hf ggml-org/Kev-4B-GGUF
```
A request contains a state and one or more questions. There are three question types:
| Type | You send | You get |
| --- | --- | --- |
| `choice` | options, with optional descriptions | the top option, plus a probability per option |
| `score` | 2 to 10 levels, lowest first | the expected level (can fall between two) |
| `noul` | a yes/no question | the probability of yes |
Send a request with your state and questions:
```
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]
}
}
}'
```
Response (values rounded):
```
{
"model": "ggml-org/Kev-4B-GGUF",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
"confidence": 0.8574
},
"angry": {
"type": "noul",
"noul": 0.8208
},
"urgency": {
"type": "score",
"score": 2.2821,
"legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
"probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
"confidence": 0.2821
}
},
"usage": {"input_tokens": 130, "output_tokens": 0}
}
```
The full reference is in the [server docs](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).
## Images
Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:
```
llama serve -hf ggml-org/OpenJev-GGUF
```
For example, to classify an uploaded document:
```
import base64
import requests
with open("document.png", "rb") as f:
image = "data:image/png;base64," + base64.b64encode(f.read()).decode()
response = requests.post("http://localhost:8080/v1/systemone", json={
"state": "A file uploaded by a customer.",
"images": [image],
"questions": {
"kind": {
"type": "choice",
"instructions": "What kind of document is this?",
"criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
},
},
})
print(response.json()["answers"]["kind"]["choice"]) # invoice
```
The `state` can also be a list of chat messages. Any `image_url` part (data URL) is read as an image, same as chat completions.
## Several models, one server
In [router mode](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp), models load on demand and you pick one per request:
```
llama serve
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'
```
`/v1/models` lists the ids. With a single model loaded, the `model` field is ignored.
## Tips
- **Try several models, of different sizes.** Small models are faster, large ones know more. The [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) compares them.
- **Describe your options.** Julia-1 routed "I was charged twice" to `shipping` with bare labels, and to `billing` (0.99) once each option had a description.
- **Pick your confidence cutoff per model.** A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
- **Batch your questions.** They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
- **Try different quantizations.** Like any GGUF, these models come in several precisions, for example `llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0`.
## What's next
New open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.
## Models mentioned in this article 5
## Collections mentioned in this article 1
More from this author
[## Using OCR models with llama.cpp
44
April 10, 2026](https://huggingface.co/blog/ggml-org/using-ocr-models-with-llama-cpp)
[## New in llama.cpp: Anthropic Messages API
48
January 19, 2026](https://huggingface.co/blog/ggml-org/anthropic-messages-api-in-llamacpp)
### Community
EditPreview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Comment
· [Sign up](https://huggingface.co/join?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp) or [log in](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp) to comment
[Upvote
14](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp)
- +2
## Models mentioned in this article 5
## Collections mentioned in this article 1 | paf1138 | 504 | 124 |
| 🟧 echo.blog ⭐ | "llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed que | ggml-org (Xuan-Son Nguyen/ngxson, Victor Mustar) | — | — |
| 🟧 hn | Rethinking LLM Serving with System One Models | matt_d | 1 | 0 |
| 🟧 hn | New in Llama.cpp: Decision Models | nreece | 3 | 0 |
| 🟧 hn | Decision 2.0: our newest decision models | thedima | 3 | 0 |
| 🟧 hn | How we built a browser agent with Jev instead of an LLM | sozal | 2 | 0 |
| 🟠 reddit | Smallest Jev-like model LocalLLaMA | TheRealREZOR | 307 | 92 |
| 🟧 hn | Show HN: Jiwo - Small decision models topping decision index leaderboard | jiwidi | 1 | 0 |
| 🟧 hn | Your Language Model Is Not-So-Secretly a System 1 Model | atbhtunnm | 1 | 1 |
| 🟧 hn | Strands Decider 2B: a small, open-source, decision model | gmays | 284 | 78 |
| 🟠 reddit | Cloudflare Clef Experience LocalLLaMA | rm-rf-rm | 0 | 17 |
| 🟧 hn | TypeSafe started the decision-model category;OpenAI and AWS did it in 3 weeks | faridovski | 5 | 2 |
| 🟠 reddit | jevman: AI decision models play Pac-Man LocalLLaMA | facethef | 271 | 79 |
| 🟠 reddit | PSA: llama.cpp PR #30100 provides a huge improvement in accuracy for Clef models on MacOS / Metal LocalLLaMA | yuicebox | 0 | 5 |
| 🟠 reddit | Auto mode plugin for the Pi coding agent that uses Kev/Laya (running locally) or Jev to classify commands LocalLLaMA | deepu105 | 9 | 4 |
| 🟠 reddit | Running decision model locally on an RTX 4090 to find out which one is the fastest LocalLLaMA | Fun-Meaning-6474 | 372 | 55 |
| 🟠 reddit | PR merged: `/v1/systemone` support in vllm LocalLLaMA | 413205 | 23 | 2 |
| 🟠 reddit | H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev) LocalLLaMA | pseudotensor1234 | 56 | 30 |
2026-10-10T03:24:26Z
vLLM merged native /v1/systemone support (reddit.post.1x21of7) and H2O-Lightning-4B claimed #1 open JevBench (reddit.post.1x1w1nv), adding a third major runtime and a new Apache-2.0 model to the ecosystem. Multi-runtime corroboration is now three-deep (llama.cpp, Ollama, vLLM); weekly model cadence holds (Clef verified, H2O new). Agent-harness adoption on the endpoint specifically remains the last unobserved clause (Pi plugin uses Jev/Kev/Laya locally; Strands Decider 2B wire-format compliance unverified). Measured heat cooled to 10.7 pts/h from 314 peak but peer percentile 89.3 and hot topic neighbourhood justify sustained attention.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x1w1nv — H2O-Lightning-4B claims #1 open JevBench score, a new Apache-2.0 decision model directly feeding the open decision-model ecosystem served by llama.cpp's /v1/systemone.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x21of7 — vLLM merging /v1/systemone support independently corroborates the decision-model endpoint pattern spreading beyond llama.cpp.
2026-10-09T07:54:19Z
grounded: converges/high — Converges with dated receipts: ggml-org, Ollama, and vLLM-side ecosystem have independently shipped the typed one-pass decision primitive Scott already runs in
2026-10-09T07:40:57Z
New independent benchmark (1x0wg85) directly tests multiple open decision models via llama.cpp's /v1/systemone endpoint; Pi coding agent plugin adopts Jev/Kev/Laya locally for command classification (harness-level decision-model use, endpoint usage unconfirmed); llama.cpp PR #30100 fixes Clef accuracy on Metal. Multi-runtime spread and ecosystem periphery continue expanding, but the hypothesis's specific confirmation clause — agent harnesses building on the /v1/systemone endpoint itself — remains unobserved; vLLM support still unverified. Current engagement rate is quiet (0 pts/h) though topic neighbourhood is hot and magnitude valve eligible.
2026-10-08T23:06:43Z
evidence attached: reddit.post.1x0wg85 — Independent user benchmark of multiple open decision models (Liquid d1, Interfaze Lev, Cloudflare Clef-Flash, Laya) via llama.cpp's /v1/systemone endpoint, directly testing the case's hypothesis about decision models becoming a standard primitive.
2026-10-08T16:09:36Z
Three new substantive evidence pieces: Pi coding agent plugin adopts Jev/Kev/Laya for command classification (harness-level decision-model use, though /v1/systemone endpoint usage unconfirmed); llama.cpp PR #30100 fixes Clef accuracy on Metal (active maintenance of the model collection); independent Pac-Man benchmark evaluates six decision models in real-time gameplay. Multi-runtime spread and ecosystem periphery continue expanding (Ollama, llama.cpp, Strands cloud-vendor entry, TinyDecide cross-platform, Jiwo, independent lab), but the hypothesis's specific confirmation clause — agent harnesses building on the /v1/systemone endpoint itself — remains unobserved; vLLM support still unverified.
2026-10-08T15:50:38Z
evidence attached: reddit.post.1x0py31 — Pi coding agent plugin adopts Jev/Kev/Laya decision models for command classification, demonstrating harness-level adoption of the decision-model primitive.
2026-10-08T15:50:38Z
evidence attached: reddit.post.1x0sj5s — llama.cpp PR #30100 fixes Clef (a decision model) accuracy on Metal, improving a model in the open decision-model collection.
2026-10-08T15:50:38Z
evidence attached: reddit.post.1x0sm1b — Independent benchmark of six decision models (Jev, Kev, Clef, GPT-6 Luna, Laya) playing Pac-Man, providing real-world evaluation of the decision-model primitive.
2026-10-08T11:53:18Z
Industry commentary video confirms OpenAI/AWS entered the decision-model category started by TypeSafe, reinforcing the category narrative, but the hypothesis's specific confirmation clause — agent harnesses building on /v1/systemone — remains unobserved. Strands Decider 2B advances the agent-stack-adoption clause as pattern adoption (first mainstream cloud-vendor entrant), though wire-format/endpoint compliance is unverified. Cloudflare Clef now verified working on the native endpoint. Measured heat has cooled to a low tail (1.2 pts/h) but Scott's briefing flag, actionability for his stack, and magnitude-valve-eligible multi-platform spread justify high heat for day-level attention.
2026-10-08T10:42:27Z
evidence attached: hn.story.50003920 — Industry commentary video arguing OpenAI/AWS entered the decision-model category started by TypeSafe; directly discusses the category llama.cpp just shipped.
2026-10-08T07:27:19Z
Cloudflare Clef model now verified working on llama.cpp's native /v1/systemone endpoint (independent user test), confirming the weekly cadence commitment; AWS-backed Strands Decider 2B continues climbing (277 pts, 78 comments, repeated velocity spikes) as the first mainstream cloud-vendor agent-stack entrant adopting the decision-model pattern — advancing the agent-stack-adoption clause, though wire-format/endpoint compliance remains unverified. Scott explicitly flagged for briefing. The hypothesis's remaining unconfirmed clause is agent harnesses building on /v1/systemone specifically.
2026-10-08T07:01:41Z
evidence attached: reddit.post.1x0jhna — User tests Cloudflare Clef model via llama.cpp's /v1/systemone decision-model endpoint, providing independent usage evidence for the newly shipped endpoint.
2026-10-07T04:38:42Z
The periphery's recruiting jumped a tier: AWS-backed Strands shipped Decider 2B, a small open decision model — the first mainstream cloud-vendor agent-stack entrant after indies (TinyDecide, Jiwo), the closest approach yet to the hypothesis's agent-stack-adoption clause, though it is a model release, not observed /v1/systemone or wire-format adoption. That new fact plus a re-warming tail (63 pts and climbing from 38, ~6.3 pts/h steady, 87.5th percentile) lifts heat low→medium; state stays at significant — the ladder tops out there and the endpoint-specific harness clause remains unobserved.
2026-10-07T03:33:13Z
evidence attached: hn.story.49987076 — AWS-backed Strands shipping a small open decision model is exactly the other-agent-stacks-adopt decision-model evidence that case tracks.
2026-10-06T18:03:01Z
The only new evidence since the promotion is a score-1 conceptual essay ('Your Language Model Is Not-So-Secretly a System 1 Model') echoing the case's thesis — no new runtime, model, harness, or adoption — so the established-ecosystem read stands unchanged and the case cools from medium to low: the live tail halved to ~2.7 pts/h and momentum is cooling. The magnitude-valve spread reading reflects the already-priced Oct 2–5 launch window plus TinyDecide's spike; current additions are comment drift, so nothing warrants hours-level attention. The case's meaning is now confirmed infrastructure to act on (migrate jevkit to the native endpoint), with harness adoption on /v1/systemone and the Cloudflare Clef cadence checkpoint as the remaining live questions.
2026-10-06T16:42:15Z
evidence attached: hn.story.49980496 — Independent technical analysis arguing LMs are System 1 models directly supports the conceptual basis of the decision-model-endpoint case.
2026-10-05T19:18:29Z
grounded: converges/high — Converges with dated receipts: ggml-org, Ollama and the vLLM-side ecosystem have independently shipped the typed one-pass decision primitive Scott already runs
2026-10-05T19:09:34Z
The format now reproduces itself below the vendor tier — a 10M embedded port (TinyDecide, browser→ESP32, quality contested) and indie fine-tunes topping the public Decision Index — which is what 'standard cheap primitive' looks like from underneath; with first-party multi-runtime support already confirmed (llama.cpp, Ollama, vLLM-side), the case is an established ecosystem, not a moving wave, so it promotes to significant. Heat lifts low→medium because the tail is live again (~13.7 pts/h steady, TinyDecide a 95.8th-percentile mover across 3 platforms) rather than the dead ~0.3 pts/h tail the last look priced — high stays reserved since the episode doesn't dominate any platform (HN posts still flopped).
2026-10-05T17:32:20Z
evidence attached: hn.story.49967455 — Independent third party releasing Qwen3.5 decision-model fine-tunes topping the public decision index is direct ecosystem-adoption evidence for decision models becoming a standard primitive.
2026-10-05T17:32:20Z
evidence attached: reddit.post.1wybi2y — A 6MB Jev-like model running from browser to ESP32, drawing real attention (79/43), is low-end spread evidence for single-pass decision models becoming a commodity primitive.
2026-10-05T16:02:56Z
First builder-adoption signal arrives: a shipped browser agent driven by Jev typed decisions instead of an LLM gives the typed-decision agent pattern its first shipped instance — but via TypeSafe's Jev rather than the local endpoint, so the confirmation clause narrows to harnesses building on /v1/systemone itself, still unobserved. State holds at accelerating because the periphery keeps adding implementations (vLLM-side Decision 2.0, this agent) even though each is thin and live attention is a cold ~0.3 pts/h tail; the magnitude-valve spread reading is residue of the Oct 2 peak and the new arrivals came in near-zero.
2026-10-05T15:30:48Z
evidence attached: hn.story.49965022 — A shipped browser agent driven by Jev typed decisions instead of an LLM is builder-adoption evidence for the typed-decision-primitive pattern that case is accelerating toward.
2026-10-04T23:26:49Z
vLLM-side 'Decision 2.0' decision models push the typed-decision primitive beyond the local ecosystem — three runtimes (Ollama, llama.cpp, vLLM) shipping first-party support inside a week is implementation-wave spread with an influential entrant, so the case accelerates on substance while attention stays a cold ~0.5 pts/h tail; the confirmation clause now narrows to its still-unseen half: agent harnesses actually building on /v1/systemone, with the Cloudflare Clef release the next weekly-cadence checkpoint.
2026-10-04T23:24:41Z
evidence attached: hn.story.49958655 — vLLM shipping its own 'Decision 2.0' models is exactly the other-runtime adoption the typed-decision-primitive hypothesis calls for.
2026-10-04T09:24:56Z
The four sensor firings resolve to long-tail vote accumulation (+60pts/+10 comments in ~36h, ~2 pts/h vs 78 peak) with zero new substance: no new runtime, no Cloudflare Clef, no harness integration, and the HN echo stayed dead (1pt/3pts, 0 comments). Meaning is unchanged from the Oct 3 adoption-watch framing; attention drops to low — the magnitude-valve eligibility is a residue of the Oct 2 multi-platform peak, while the live curve is a single-Reddit tail.
2026-10-03T08:42:42Z
The episode's attention has peaked and is cooling (Reddit decelerated ~78→8 pts/h, the new HN arrivals flopped at 3pts/0 comments), so this stops being a breaking-attention story and becomes a corroborated adoption-watch case; nothing new about the hypothesis itself — next meaning-changers are Cloudflare Clef, the first weekly-release checkpoint, and any agent harness actually building on /v1/systemone.
2026-10-03T08:24:40Z
evidence attached: hn.story.49942092 — Same ggml-org decision-models announcement spreading to HN front page, adding community spread to the open case.
2026-10-02T20:22:48Z
Material change: the previously unverified first-party claim is now confirmed by ggml-org's own retrieved HF blog (Oct 2, 2026, Nguyen/Mustar) documenting the /v1/systemone endpoint, PR #29818, the exact Julia-1→OpenJev five-model collection, weekly release cadence, and Cloudflare Clef as next — retiring the grounding caveat. Case moves seed→corroborated on the format's multi-runtime spread (Ollama native, third-party llama-server adapter, independent lab writeup); top-decile multi-platform velocity (94.6th percentile) with the periphery still adding outlets justifies high attention despite steady momentum.
2026-10-02T19:29:58Z
evidence attached: hn.story.49936590 — Independent AI-lab writeup arguing System One models as a serving primitive is third-party engagement with exactly the decision-model-serving thesis this case tracks.
2026-10-02T15:36:47Z
grounded: converges/high — Scott's own wikis already carry System One typed decisions in production (dev:project.jev warm-start with spam-gate and mail-front-door generations, dev:concept
2026-10-02T15:27:02Z
case created — First-party, already-shipped infrastructure (merged PR #29818, five-model collection live) making typed one-pass decisions a native primitive of the dominant local runtime — a distinct runtime-adoption episode from the existing Jev model-quality, vendor-API, and third-party-adapter cases.