2026-10-11 16:37 UTC

ggml-org claims llama.cpp's newly shipped /v1/systemone decision-model endpoint — serving an open collection from 144M Julia-1 to vision-capable 27B OpenJev with 'new open decision models every week' — makes single-forward-pass typed decisions (routing, moderation, compaction checks, agent next-action) a standard cheap primitive of the dominant local runtime; adoption by local agent stacks and other runtimes following the System One format confirms it, stalled uptake refutes it.

state: significantheat: mediumuncertainty: lowconvergesscott: highlocal-inference agent-harnesses agent-orchestration llamacppggml-orgXuan-Son NguyenVictor MustarTypeSafe
Surfaced 2026-10-02T20:24:50Z — "llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed que — Material change: the previously unverified first-party claim is now confirmed by ggml-org's own retrieved HF blog (Oct 2, 2026, Nguyen/Mustar) documenting the /v1/systemone endpoint, PR #29818, the exact Julia-1→OpenJev five-model collection, weekly release cadence, and Cloudflare Clef as next — retiring the grounding caveat. Case moves seed→corroborated on the format's multi-runtime spread (Ollama native, third-party llama-server adapter, independent lab writeup); top-decile multi-platform velocity (94.6th percentile) with the periphery still adding outlets justifies high attention despite steady momentum.

What is this?

ggml-org (Xuan-Son Nguyen, Victor Mustar) shipped a native `/v1/systemone` endpoint in llama.cpp server (PR #29818, announced Oct 2 2026) that serves an open collection of five decision models — Julia-1 144M, Laya 421M, Kev-4B, lev 4B, and vision-capable OpenJev 27B — returning per-option probabilities in a single forward pass with zero generated tokens. The System One wire format (state + typed questions: choice/score/noul) originated from TypeSafe's Jev; llama.cpp's implementation uses existing GBNF grammar, n_probs, and prompt-cache primitives. Multi-runtime adoption followed within days: Ollama native `/v1/systemone` (v0.35, Sept 29), vLLM-side 'Decision 2.0' models (Oct 4-5), and AWS-backed Strands shipping Decider 2B as the first mainstream cloud-vendor agent-stack entrant. Independent ecosystem signals include TinyDecide 10M (browser→ESP32), Jiwo fine-tunes topping the Decision Index, a shipped browser agent using Jev, Pi coding agent plugin using Jev/Kev/Laya for command classification, and multiple independent benchmarks via the llama.cpp endpoint. The weekly open-model cadence commitment is being honored: Cloudflare Clef verified working on the endpoint via bartowski quant (Oct 8).

Why it matters to Scott

Converges with dated receipts: ggml-org, Ollama, and vLLM-side ecosystem have independently shipped the typed one-pass decision primitive Scott already runs in production (dev:project.jev warm-start generations, dev:concept.cheap-model-front-door). The native /v1/systemone endpoint on his gamepc replaces the third-party Verdict adapter for jevkit and the cheap-model front door; per-model confidence cutoffs can be tuned on his examples; multi-runtime support keeps endpoint choice portable.
dev:project.jevdev:concept.cheap-model-front-doordev:technology.typesafe-jevip:concept.system-one-typed-decisionsdev:project.gamepcip:framework.open-weights-sovereigntydev:concept.agent-orchestrationip:concept.micro-judgement-patternradar:typesafe-jev-structured-decisionsradar:verdict-local-jev-compatible-decisionsradar:system-one-harness-typed-action-loopradar:jevman-pacman-decision-model-benchmarkradar:strands-harness-releaseradar:open-weight-inference-economics
queries asked of Scott's wikis
  • dev:project.jev warm-start generations spam-gate mail-front-door
  • dev:concept.cheap-model-front-door local-first routing escalation
  • ip:concept.system-one-typed-decisions one-forward-pass primitive
  • dev:project.agent-harnesses decision-model integration patterns
  • ip:framework.open-weights-sovereignty local inference economics
  • dev:concept.agent-orchestration routing moderation compaction next-action

Measured heat

now 0 pts/hpeak 315 pts/hcomments 0/hpeers p25momentum: steady3 platformsage 242h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-01 14:00⭐ origin echo-reconstructed"llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed que
ggml-org (Xuan-Son Nguyen/ngxson, Victor Mustar) on blog (echo) · attributed from reddit.post.1wvv6im
—
10-02 14:25first on r/LocalLLaMA · published · +24.4hNew in llama.cpp: Decision Models
paf1138
—
10-02 18:02first on hacker news · published · +28.1hRethinking LLM Serving with System One Models
matt_d
—
10-02 14:25amplified on r/LocalLLaMAreddit.post.1wvv6im
paf1138
peak 504 · 124 comments · 24% of case engagement
10-02 18:02amplified on hacker newshn.story.49936590
matt_d
peak 1 · 0 comments · 0% of case engagement
10-03 07:27amplified on hacker newshn.story.49942092
nreece
peak 3 · 0 comments · 0% of case engagement
10-04 22:36amplified on hacker newshn.story.49958655
thedima
peak 3 · 0 comments · 0% of case engagement
10-05 14:00amplified on hacker newshn.story.49965022
sozal
peak 2 · 0 comments · 0% of case engagement
10-05 15:28amplified on r/LocalLLaMAreddit.post.1wybi2y
TheRealREZOR
peak 307 · 92 comments · 15% of case engagement
11 more amplifiers in ainews.case_chain
10-02 15:20our radar first saw it · +25.3hdiscovery anchor: reddit.post.1wvv6im—
10-02 20:22reached heat=high · +30.4h · via ledger——
pace: p93 vs 1188 stories at the 168h mark (now 242h old) — ahead of minab-maven-target-verification-failure (1.0x), behind heretic-model-unrestriction-tooling (1.0x)

Evidence (18) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNew in llama.cpp: Decision Models
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-10-02T15:25:30.921317+00:00

[Back to Articles](https://huggingface.co/blog)

# New in llama.cpp: Decision Models

[Community Article](https://huggingface.co/blog/community)

Published
October 2, 2026

[Upvote

14](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp)  







 - +8

[Xuan-Son Nguyen's avatar](https://huggingface.co/ngxson) 

[Xuan-Son Nguyen

ngxson 

Follow](https://huggingface.co/ngxson)

[ggml-org's avatar](https://huggingface.co/ggml-org "ggml-org") [ggml-org](https://huggingface.co/ggml-org)

[Victor Mustar's avatar](https://huggingface.co/victor) 

[Victor Mustar

victor 

Follow](https://huggingface.co/victor)

[ggml-org's avatar](https://huggingface.co/ggml-org "ggml-org") [ggml-org](https://huggingface.co/ggml-org)

 

llama.cpp server now supports decision models through the `/v1/systemone` endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.

The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in [PR #29818](https://github.com/ggml-org/llama.cpp/pull/29818).

> **What is a decision model?** A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.

## Supported models

| Model | Size | Based on | Languages | Images | License | Speed\* |
| --- | --- | --- | --- | --- | --- | --- |
| [Julia-1](https://huggingface.co/ggml-org/Julia-1-GGUF) | 144M | mmBERT-small | 50+ | no | Apache 2.0 | 3 ms |
| [Laya](https://huggingface.co/ggml-org/Laya-GGUF) | 421M | ModernBERT-large | English | no | Apache 2.0 | 5 ms |
| [Kev-4B](https://huggingface.co/ggml-org/Kev-4B-GGUF) | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | 12 ms |
| [lev](https://huggingface.co/ggml-org/lev-GGUF) | 4B | Qwen3.5-4B | English | no | Apache 2.0 | 36 ms |
| [OpenJev](https://huggingface.co/ggml-org/OpenJev-GGUF) | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | 43 ms |

\*Median time to answer one question, on one NVIDIA RTX PRO 6000.

Find these models in the [Decision models collection](https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769), with more coming. The community [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) shows how they compare.

## Quick start

Get the latest llama.cpp from [llama.app](https://llama.app) (or run `llama update`), then start a model:

```
llama serve -hf ggml-org/Kev-4B-GGUF
```

A request contains a state and one or more questions. There are three question types:

| Type | You send | You get |
| --- | --- | --- |
| `choice` | options, with optional descriptions | the top option, plus a probability per option |
| `score` | 2 to 10 levels, lowest first | the expected level (can fall between two) |
| `noul` | a yes/no question | the probability of yes |

Send a request with your state and questions:

```
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'
```

Response (values rounded):

```
{
  "model": "ggml-org/Kev-4B-GGUF",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
      "confidence": 0.8574
    },
    "angry": {
      "type": "noul",
      "noul": 0.8208
    },
    "urgency": {
      "type": "score",
      "score": 2.2821,
      "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
      "probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
      "confidence": 0.2821
    }
  },
  "usage": {"input_tokens": 130, "output_tokens": 0}
}
```

The full reference is in the [server docs](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).

## Images

Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:

```
llama serve -hf ggml-org/OpenJev-GGUF
```

For example, to classify an uploaded document:

```
import base64
import requests

with open("document.png", "rb") as f:
    image = "data:image/png;base64," + base64.b64encode(f.read()).decode()

response = requests.post("http://localhost:8080/v1/systemone", json={
    "state": "A file uploaded by a customer.",
    "images": [image],
    "questions": {
        "kind": {
            "type": "choice",
            "instructions": "What kind of document is this?",
            "criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
        },
    },
})

print(response.json()["answers"]["kind"]["choice"])  # invoice
```

The `state` can also be a list of chat messages. Any `image_url` part (data URL) is read as an image, same as chat completions.

## Several models, one server

In [router mode](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp), models load on demand and you pick one per request:

```
llama serve
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'
```

`/v1/models` lists the ids. With a single model loaded, the `model` field is ignored.

## Tips

- **Try several models, of different sizes.** Small models are faster, large ones know more. The [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) compares them.
- **Describe your options.** Julia-1 routed "I was charged twice" to `shipping` with bare labels, and to `billing` (0.99) once each option had a description.
- **Pick your confidence cutoff per model.** A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
- **Batch your questions.** They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
- **Try different quantizations.** Like any GGUF, these models come in several precisions, for example `llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0`.

## What's next

New open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.

## Models mentioned in this article 5

    

## Collections mentioned in this article 1

 

More from this author

[## Using OCR models with llama.cpp

44

 April 10, 2026](https://huggingface.co/blog/ggml-org/using-ocr-models-with-llama-cpp)

[## New in llama.cpp: Anthropic Messages API

48

 January 19, 2026](https://huggingface.co/blog/ggml-org/anthropic-messages-api-in-llamacpp)

 

### Community

EditPreview

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Comment 

· [Sign up](https://huggingface.co/join?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp) or [log in](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp) to comment

[Upvote

14](https://huggingface.co/login?next=%2Fblog%2Fggml-org%2Fdecision-models-in-llamacpp)  













 - +2

## Models mentioned in this article 5

    

## Collections mentioned in this article 1
paf1138504124
🟧 echo.blog ⭐"llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed queggml-org (Xuan-Son Nguyen/ngxson, Victor Mustar)——
🟧 hnRethinking LLM Serving with System One Modelsmatt_d10
🟧 hnNew in Llama.cpp: Decision Modelsnreece30
🟧 hnDecision 2.0: our newest decision modelsthedima30
🟧 hnHow we built a browser agent with Jev instead of an LLMsozal20
🟠 redditSmallest Jev-like model
LocalLLaMA
TheRealREZOR30792
🟧 hnShow HN: Jiwo - Small decision models topping decision index leaderboardjiwidi10
🟧 hnYour Language Model Is Not-So-Secretly a System 1 Modelatbhtunnm11
🟧 hnStrands Decider 2B: a small, open-source, decision modelgmays28478
🟠 redditCloudflare Clef Experience
LocalLLaMA
rm-rf-rm017
🟧 hnTypeSafe started the decision-model category;OpenAI and AWS did it in 3 weeksfaridovski52
🟠 redditjevman: AI decision models play Pac-Man
LocalLLaMA
facethef27179
🟠 redditPSA: llama.cpp PR #30100 provides a huge improvement in accuracy for Clef models on MacOS / Metal
LocalLLaMA
yuicebox05
🟠 redditAuto mode plugin for the Pi coding agent that uses Kev/Laya (running locally) or Jev to classify commands
LocalLLaMA
deepu10594
🟠 redditRunning decision model locally on an RTX 4090 to find out which one is the fastest
LocalLLaMA
Fun-Meaning-647437255
🟠 redditPR merged: `/v1/systemone` support in vllm
LocalLLaMA
413205232
🟠 redditH2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev)
LocalLLaMA
pseudotensor12345630

Interpretation history

Decision trace