2026-10-11 16:38 UTC

System One Lite's maintainer claims its released MLX proof of concept extracts typed answer distributions from stock local-model logits without text generation or parsing, enabling closed-set software decisions without fine-tuning.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference structured-output typed-decisionssnellingio

What is this?

System One Lite is snellingio's MIT-licensed, released MLX proof of concept (github.com/snellingio/system-one, not directly returned in these snippets) that makes stock local Qwen3 models answer typed questions without generating text: the model runs to a designated answer slot, logits are restricted to valid answer-token codes, and the softmaxed result becomes a choice/score distribution โ€” no decoding, no parsing, no fine-tuning. The supplied results show this is no longer an isolated trick but one node in a small, visibly forming ecosystem of generation-free 'System One'-style decision models: TypeSafe AI's commercial Jev is the established namesake, Laya-MLX (mizorewww) is an Apple Silicon runtime using a bidirectional encoder for one-pass classifications and scores, iapp's OpenThai-SystemOne-MLX-mxfp4 on Hugging Face claims one-forward-pass typed answers with 'calibrated probabilities', an open 395M 'Von' model exists, and a r/LocalLLaMA user explicitly evaluates 'local replacements for TypeSafe Jev' โ€” finding none match the original, at least for their use. Caveats: the snippet evidence is thin โ€” OpenThai's calibration claim is an unvalidated model-card statement, Laya's encoder technique differs technically from System One Lite's stock-causal-model logit scoring, the Von assessment is a single user's report, and nothing supplied independently validates System One Lite's own accuracy, calibration, or option-order robustness.

Why it matters to Scott

Converges with the Judgment Join / answers-not-content separation Scott already productionised โ€” scoring declared answer-token codes instead of generate-and-repair is a dated third-party receipt for his typed-disposition doctrine and a concrete mechanism for the cheap-model front door. It bears on live assets: an MIT, generation-free Jev-style engine runs on the Mac-mini MLX stack he already operates and could back up or substitute for his TypeSafe Jev deployments (mail front door, Venture World stage manager) โ€” but the README's own uncalibrated-probabilities admission, zero independent validation, r/LocalLLaMA's finding that local clones trail the original, and his calm-eval caveat (a logit-restricted softmax cannot decline, so 'confidence' is relative preference, not correctness) keep this at watch-and-bounded-eval, not adopt.
ip:framework.judgment-joindev:concept.answers-not-contentdev:concept.cheap-model-front-doordev:technology.typesafe-jevdev:project.jevdev:technology.mlxip:concept.calm-evalradar:typesafe-jev-structured-decisionsradar:verdict-local-jev-compatible-decisionsradar:blink-embedded-typed-decisionsradar:intern-decision-one-pass-decisionsradar:concept.typed-decisionsradar:concept.mlx
queries asked of Scott's wikis
  • judgment join typed AI dispositions versus deterministic mutations
  • scoring declared answers instead of generate-and-repair structured extraction
  • MLX Apple silicon local inference projects and stack
  • model confidence calibration thresholding labeled eval data
  • closed-set decision routing in agent harnesses
  • logit readout constrained decoding without generation

Measured heat

now 0 pts/hpeak 17 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 601h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-16 15:23 (minted)โญ origin echo-reconstructedSystem One Lite scores declared answer-token codes using stock MLX models and returns typed distributions with zero generated tokens; it exp
snellingio on github (echo) ยท attributed from hn.story.49727643 ยท published time unknown
โ€”
09-16 14:31first on hacker news ยท published ยท lag ?System One Lite โ€“ typed decisions from a local LLM, with no generated tokens
samsnelling
โ€”
10-05 20:01first on r/LocalLLaMA ยท published ยท lag ?nokia-applied-research/AnyJev: Turn any LLM into a Jev-style decision model
Atagor
โ€”
09-16 14:31amplified on hacker newshn.story.49727643
samsnelling
peak 2 ยท 0 comments ยท 5% of case engagement
10-05 20:01amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyinzn
Atagor
peak 55 ยท 6 comments ยท 89% of case engagement
10-06 12:38amplified on r/LocalLLaMAreddit.post.1wz1ime
piotr1215
peak 1 ยท 3 comments ยท 6% of case engagement
09-16 15:20our radar first saw it ยท lag ?discovery anchor: hn.story.49727643โ€”
pace: p23 vs 1032 stories at the 336h mark (now 601h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (4) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnSystem One Lite โ€“ typed decisions from a local LLM, with no generated tokens
Retrieved article excerpt

Open article ยท Retrieved 2026-09-16T15:22:40.638365+00:00

# System One Lite

A tiny project that turns a normal local LLM into a typed decision engine. It
needs no fine-tuning, text generation, or parser. [Read the docs](https://github.com/snellingio/system-one/blob/main/docs/index.md).

## Stop asking language models to write. Start making them decide.

Language models are brilliant at producing text. Software does not want text.
It wants a route, a score, a yes or no, and an honest signal when the answer is
unclear.

So why are we still asking models to write tiny essays? We parse those essays
back into data, validate the data, retry failures, and hope nothing goes off the
rails.

**System One Lite deletes the essay.**

Send one unstructured state plus up to 64 typed questions. A local model scores
only the answers you allow and returns a probability distribution for every
question.

**Unstructured state in. Typed probabilities out. Zero generated tokens.**

No free-form response. No JSON repair loop. No invented option that your code
has never heard of.

This is an independent proof of concept.

It uses the System One Models interface.
It is not a new foundation model. It runs a stock open-weight model with MLX.
The project tests simpler AI software where the model can only decide.

## The old stack is absurd

|  | A normal LLM feature | System One Lite |
| --- | --- | --- |
| Input | Unstructured text | Unstructured text or JSON |
| Output | A generated string | A typed answer over declared options |
| Uncertainty | A guess written in prose | The full probability distribution |
| Validation | Parse, validate, retry | Constrained by construction |
| Output tokens | One token at a time | **Zero** |
| Failure mode | Malformed data or invented values | A valid answer that may still be wrong |
| Deployment | Usually a hosted model | A fixed local model on Apple silicon |

The final row matters. System One Lite does not make a small model infallible.
It makes the limit between the model and your code brutally clear. The
model can choose the wrong declared answer. It cannot create a new one.

That is the difference between asking AI to behave like an API and giving it
an interface it cannot break.

## Three primitives. A ridiculous number of decisions.

| Type | Ask it to | Get back |
| --- | --- | --- |
| [Choice](https://github.com/snellingio/system-one/blob/main/docs/primitives/choice.md) | Pick from a closed set | Winner, probabilities, and confidence |
| [Score](https://github.com/snellingio/system-one/blob/main/docs/primitives/score.md) | Judge a position on an ordered scale | Weighted score, level probabilities, and confidence |
| [Noul](https://github.com/snellingio/system-one/blob/main/docs/primitives/noul.md) | Make a yes or no judgment | Probability of yes |

Route support tickets. Rank leads. Gate a workflow. Score risk. Flag content.
Decide whether a human needs to look. Combine several small judgments into a
larger rule that stays in ordinary code.

Each question is independent. One answer cannot leak into the next. Your code,
not a hidden chain of thought, decides what happens after the probabilities
arrive.

## Watch it decide

System One Lite needs an Apple silicon Mac, Python 3.12 or newer, and
[`uv`](https://docs.astral.sh/uv/). Start the server:

```
cd server
uv sync
uv run uvicorn system_one_lite.api:app --port 8010
```

The first start loads `mlx-community/Qwen3-1.7B-4bit` and compiles the
Metal kernels. Then send a request:

```
curl -s http://127.0.0.1:8010/evaluate \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "state": "My order was due Friday, but it is still in transit.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this message?",
      "criteria": {
        "deliveries": "Late, missing, or damaged orders",
        "billing": "Charges, refunds, or payment methods",
        "account": "Login, profile, or app problems"
      }
    },
    "needs_reply": {
      "type": "noul",
      "instructions": "Does the customer need a reply?"
    }
  }
}
EOF
```

One request comes back ready for code:

```
{
  "model": "mlx-community/Qwen3-1.7B-4bit",
  "answers": {
    "team": {
      "type": "choice",
      "choice": "deliveries",
      "probabilities": {
        "deliveries": 0.71,
        "billing": 0.18,
        "account": 0.11
      },
      "confidence": 0.565
    },
    "needs_reply": {
      "type": "noul",
      "noul": 0.88
    }
  },
  "usage": {
    "input_tokens": 94,
    "output_tokens": 0
  }
}
```

The numbers show the response shape. Exact values depend on the input and
model. The shape does not.

For a longer example with all three question types, open the
[quickstart](https://github.com/snellingio/system-one/blob/main/docs/quickstart.md).

## Pick the model

The 1.7B model is the `default` profile. The 4B Instruct model is the
`larger` profile. Download either model before a run:

```
uv run python -m tools.download_model default
uv run python -m tools.download_model larger
```

The demo, benchmark, and eval tools accept either profile. For example:

```
uv run python -m tools.evals --model larger --limit 20
```

The server uses `default` unless `SYSTEM_ONE_MODEL` selects another profile:

```
SYSTEM_ONE_MODEL=larger uv run uvicorn system_one_lite.api:app --port 8010
```

Both model repositories are pinned to exact commits and have checked-in
answer-code registries. Change models only after you run the same accuracy and
option-order checks on both.

## The trick is almost offensively simple

For every question, the server:

1. Writes the state, question, and allowed answers into a prompt.
2. Assigns each answer a token code such as `A`, `B`, or `C`.
3. Runs the model up to the answer slot without decoding any text.
4. Throws away every logit except the valid answer codes.
5. Applies softmax and maps the probabilities back to your labels.

The model never gets the chance to ramble. It reaches the exact point where
an answer must appear, and System One Lite reads the scores directly.

Choice returns the winning label and the full distribution. Score returns the
probability-weighted level. Noul returns the probability of `yes`. Every extra
question gets its own full prompt, so questions cannot affect one another.

Read [How it works](https://github.com/snellingio/system-one/blob/main/docs/how-it-works.md) for token alignment, option limits,
and the reason this safe path uses one model pass per question.

## Extraordinary claims, meet a local eval

There is no benchmark confetti here. The repository has an eval runner for
JSONL files that use the documented dataset envelope:

```
cd server
uv run python -m tools.evals --datasets /path/to/jsonl-directory --limit 20
```

The report shows accuracy by question type. It also rotates Choice options and
checks whether changing their order changes the winner. The public dataset is
the next release step and is not in Git yet. Local files under `datasets/` are
ignored, so they cannot be published by accident.

Better yet, add examples from your own traffic. A decision system earns trust
on the states it will actually see, not on a launch graphic.

## Use it from code

The repository includes two local SDKs:

- [Python](https://github.com/snellingio/system-one/blob/main/sdks/python/README.md): sync and async clients with no runtime
  dependencies. Python 3.9 or newer.
- [JavaScript](https://github.com/snellingio/system-one/blob/main/sdks/javascript/README.md): a typed client for Node 22.18 or
  newer.

Both clients use `http://127.0.0.1:8010` by default. Set `SYSTEM_BASE_URL` to
change it.

## The part most launch posts bury

System One Lite is an experiment, not a production decision service.

- The default 1.7B model is small. It will not match a frontier model on hard
  judgments.
- The returned probabilities are model scores. They are **not calibrated odds
  of being correct**.
- Synthetic eval data does not stand in for real production traffic.
- Every question repeats the state and runs separately. Cost grows with the
  number and length of questions.
- The server handles one inference request at a time and returns `503` while
  the engine is busy.
- The current server requires Apple silicon because it uses MLX.

Use confidence to route uncertain cases. Set thresholds from labeled data that
matches your traffic. Read [Confidence](https://github.com/snellingio/system-one/blob/main/docs/confidence.md) before you let a
score trigger anything expensive, sensitive, or hard to undo.

## Run every check

```
cd server
uv run ruff check src tools tests
uv run ruff format --check src tools tests
uv run pytest

cd ../sdks/python
uv run --with pytest pytest -q

cd ../javascript
npm ci
npm run typecheck
npm test
```

## Project map

| Path | What is inside |
| --- | --- |
| [`server/`](https://github.com/snellingio/system-one/blob/main/server) | FastAPI service, MLX engine, evals, and tests |
| [`sdks/python/`](https://github.com/snellingio/system-one/blob/main/sdks/python) | Python client |
| [`sdks/javascript/`](https://github.com/snellingio/system-one/blob/main/sdks/javascript) | TypeScript client |
| [`docs/`](https://github.com/snellingio/system-one/blob/main/docs) | Guides and API reference |

Start with the [introduction](https://github.com/snellingio/system-one/blob/main/docs/introduction.md). Then read the
[API reference](https://github.com/snellingio/system-one/blob/main/docs/api.md) for the complete request shape, limits, and error
responses.

## License

[MIT](https://github.com/snellingio/system-one/blob/main/LICENSE). The supported MLX model repositories use Apache 2.0.
samsnelling20
๐ŸŸง echo.github โญSystem One Lite scores declared answer-token codes using stock MLX models and returns typed distributions with zero generated tokens; it expsnellingioโ€”โ€”
๐ŸŸ  redditnokia-applied-research/AnyJev: Turn any LLM into a Jev-style decision model
LocalLLaMA
Atagor556
๐ŸŸ  redditclassif: shell scripts that branch on meaning, read from one token's logprobs on a local 12B
LocalLLaMA
piotr121513

Interpretation history

Decision trace