2026-10-11 16:38 UTC

Bespoke Labs claims its released Nimble stack β€” 2,676 curated training examples, the Bespoke-Nimble-9B checkpoint, and a full training/serving recipe β€” shows a one-day LoRA of Qwen3.5-9B can deliver Jev-style typed decisions scoring 90.1% of reference labels versus 93.2% for TypeSafe's proprietary Jev 1.13.0, making Jev-class judgment reproducible on open weights; third-party adoption or replication of the model and recipe, or its fading into a demo, resolves whether open Jev foundations become a standard agent-harness component.

state: significantheat: lowuncertainty: mediumconvergesscott: highopen-model-releases agent-harnesses local-inferenceBespoke LabsTypeSafeEdgar Dyck
Surfaced 2026-10-01T19:45:44Z β€” Repo titled 'Data, Model, Recipe for an open Jev': serving reads the prompt once and scores one answer token per question; training fine-tun β€” Cloudflare's Clef β€” open-weights Jev-class post-trains of Qwen3.8-27B and Qwen3.5-9B (clef-flash) β€” is the case's first influential infrastructure entrant and the closest thing yet to the adoption signal the last look set as its bar: standardization of open Jev-class typed decisions is now being driven at vendor level, not by one-off copies. But the entry is independent of Nimble (no recipe reuse), and the EU 'System One' service landed as a dud (0.07 ratio), so the heat rise is attributable to a single strong artifact rather than broadening periphery; Nimble-specific adoption remains absent and the Jev calibration edge is unclosed.

What is this?

Bespoke Labs released Nimble, an open (Apache-2.0) replication of TypeSafe AI's proprietary Jev β€” the 'System One' decision model that returns typed choices, booleans, and rubric scores with probabilities instead of generated text. Nimble is a ~165 MiB LoRA adapter on Qwen3.5-9B trained on ~2,676 contrastively curated examples (each flips one fact so the correct answer flips); serving reads the prompt once and scores one answer token per field off the logits, exposed through the same System One request format Jev uses. On its 324-example holdout Bespoke reports 90.1% reference-label agreement vs 93.2% for Jev 1.13.0 and 66.4% for base Qwen3.5-9B, and the release ships the dataset, adapter, prompt builder, and reference inference code. Caveat: the supplied snippets carry only Bespoke's self-reported numbers plus generic Jev explainers (Flowtivity, StackAI, LangChain's harness guide) β€” no independent verification, no license terms for the code/data, and no third-party adoption of Nimble itself appear in them.

Why it matters to Scott

Converges with his curation-is-the-moat argument (ip:concept.capability-symmetry, ip:source.the-moat-is-the-memory-ebook): a one-day LoRA on 2,676 contrastive examples landing ~2 points off proprietary Jev β€” then Cloudflare, Perplexity and llama.cpp mainline independently shipping the same open decision-model layer β€” is a dated receipt that weights aren't the moat. It bears directly on dev:technology.typesafe-jev: the single-vendor lock-in behind his three production front doors is gone, with candidate open checkpoints now behind a mainline /v1/systemone API (swap is a config change, none forced), and the recorded many-fast-low-context / high-context boundary condition feeds ip:concept.placement-judgment.
dev:technology.typesafe-jevip:concept.capability-symmetryip:source.the-moat-is-the-memory-ebookip:concept.model-perishabilitydev:concept.cheap-model-front-doorip:concept.placement-judgmentradar:typesafe-jev-structured-decisionsradar:llamacpp-decision-modelsradar:concept.typed-decisionsradar:verdict-local-jev-compatible-decisionsradar:system-one-lite-typed-decisionsradar:intern-decision-one-pass-decisions
queries asked of Scott's wikis
  • curation as moat β€” small curated/contrastive datasets beating scale
  • open-weights replication of closed proprietary model capabilities
  • agent harness decision routing β€” fast small model vs frontier call
  • LoRA adapter training and local serving stack
  • typed decisions / structured output layer in agent codebases
  • local inference economics β€” latency and cost thresholds for swapping a hosted model

Measured heat

now 3 pts/hpeak 202 pts/hcomments 1/hpeers p88momentum: cooling3 platformsage 266h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-30 14:27 (minted)⭐ origin echo-reconstructedRepo titled 'Data, Model, Recipe for an open Jev': serving reads the prompt once and scores one answer token per question; training fine-tun
Bespoke Labs (bespokelabsai) on github (echo) Β· attributed from hn.story.49909001 Β· published time unknown
β€”
09-30 13:53first on hacker news Β· published Β· lag ?Nimble: Data, Model, Recipe for an Open Jev (From Bespoke Labs)
michael-sumner
β€”
09-30 17:13first on r/LocalLLaMA Β· published Β· lag ?Gliner2.5-Decide (Jev style model)
parepeg
β€”
10-02 07:44first on r/singularity Β· published Β· lag ?I benchmarked Jev 1.13 against 4 local LLMs on RTX5070 12GB - amazing.
Storge2
β€”
10-10 20:52first on r/artificial Β· published Β· lag ?Non-text AI model maker Jev valued at $7.5B just weeks after launch
lulzxdxdxd
β€”
09-30 13:53amplified on hacker newshn.story.49909001
michael-sumner
peak 1 Β· 0 comments Β· 0% of case engagement
09-30 17:13amplified on r/LocalLLaMAreddit.post.1wuaq3l
parepeg
peak 15 Β· 7 comments Β· 2% of case engagement
09-30 17:37amplified on hacker newshn.story.49911997
jeff_ciesielski
peak 2 Β· 1 comments Β· 0% of case engagement
10-01 02:12amplified on hacker newshn.story.49916865
chrismungall
peak 12 Β· 1 comments Β· 2% of case engagement
10-01 15:44amplified on hacker newshn.story.49923223
mrkn1
peak 9 Β· 2 comments Β· 2% of case engagement
10-01 17:04amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wv4zzi
paf1138
peak 406 Β· 116 comments Β· 44% of case engagement
19 more amplifiers in ainews.case_chain
09-30 14:21our radar first saw it Β· lag ?discovery anchor: hn.story.49909001β€”
10-01 19:40reached heat=high Β· lag ? Β· via ledgerβ€”β€”
pace: p90 vs 1188 stories at the 168h mark (now 266h old) β€” ahead of c5r-astra-research-facility (1.0x), behind amazon-blocks-meta-muse-shopping (1.0x)

Evidence (27) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnNimble: Data, Model, Recipe for an Open Jev (From Bespoke Labs)
Retrieved article excerpt

Open article Β· Retrieved 2026-09-30T14:25:27.024691+00:00

# Bespoke Nimble

**Data, Model, Recipe for an open Jev**

[Model](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B) Β· [Updates](https://github.com/bespokelabsai/nimble#updates) Β· [Capabilities](https://github.com/bespokelabsai/nimble#capabilities) Β· [Quickstart](https://github.com/bespokelabsai/nimble#quickstart) Β· [Methodology](https://github.com/bespokelabsai/nimble#methodology) Β· [Documentation and development](https://github.com/bespokelabsai/nimble#documentation-and-development) Β· [Citation](https://github.com/bespokelabsai/nimble#citation)

[Introducing Bespoke Nimble. Serving reads the prompt once and then scores one answer token per question. Data curation changes one fact so that the correct answer flips. Training fine-tunes Qwen3.5-9B with LoRA on the answer tokens only. On 324 held-out examples, Bespoke-Nimble-9B matches 90.1% of the reference labels, compared with 66.4% for its base model and 93.2% for Jev 1.13.0.](https://github.com/bespokelabsai/nimble/blob/main/assets/diagrams/nimble-infographic.svg)

Nimble takes some text and a schema, and makes typed decisions about the text.
The schema is the list of questions to answer. Each question is either a choice
from a list that you give or a true or false question. For each question, Nimble
returns the answer it picked and the probability of each allowed answer.

Nimble makes each decision in one step and does not write out any reasoning
first, so it is fast (blazing fast!). Nimble is
inspired by the System One approach of
[TypeSafe's Jev](https://docs.typesafe.ai/primitives/choice). In this repository,
we share our recipe for training such a model.

Note that we did not distill from Jev. The point of the repository is to show how to curate data, how to train, and to serve such a model, and encourage more research!

You can run [Bespoke-Nimble-9B](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B)
on a Mac with Apple Silicon or on a machine with an NVIDIA GPU.

## Updates

- September 24, 2026: [Bespoke-Nimble-9B](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B) now contains the latest checkpoint, with an 8,192-token context and up to 255 choices per field. Its default is T=1.0; the original release is preserved under the `original-2676` tag, and the separate v2 repository is unchanged. Update this checkout before loading the new release.
- On September 22, 2026, we fitted a temperature for Bespoke-Nimble-9B. With
  this temperature, the probabilities better match how often the answers are
  right. The model picks the same answers as before. Noul probabilities and
  Score values do change, so if you compare them with a threshold, test the
  threshold again. See [Probability temperature](https://github.com/bespokelabsai/nimble#probability-temperature) and
  [PR #7](https://github.com/bespokelabsai/nimble/pull/7).
- On September 20, 2026, we published the 2,676 training examples and the 324
  held-out examples for Bespoke-Nimble-9B. We had left them out of the first
  release by mistake. See the [dataset guide](https://github.com/bespokelabsai/nimble/blob/main/docs/DATASET.md) and
  [PR #5](https://github.com/bespokelabsai/nimble/pull/5).
- On September 19, 2026, we raised the prompt limit of the hosted API to 8,192
  tokens for each question. The model was trained on prompts of up to 2,048
  tokens, so shorter prompts are better tested. See the
  [SGLang deployment guide](https://github.com/bespokelabsai/nimble/blob/main/docs/MODAL_SERVING.md) and
  [PR #4](https://github.com/bespokelabsai/nimble/pull/4).
- On September 18, 2026, [Edgar Dyck](https://github.com/eddited17) added a
  public benchmark suite. With it, you can run Bespoke-Nimble-9B and Jev on the
  same records from 13 public subsets with human labels. The
  [public benchmarks guide](https://github.com/bespokelabsai/nimble/blob/main/docs/PUBLIC_BENCHMARKS.md) has the steps and the
  results. See [PR #2](https://github.com/bespokelabsai/nimble/pull/2).

## Capabilities

We built Nimble in one day, so expect some rough edges. What Nimble can do comes
from two sources: the first is the base model, Qwen3.5-9B, the second is our
training data, which we curated for a few specific domains.

### What you can build

| Task | You define | You get back |
| --- | --- | --- |
| Route a request | The destinations and when each one applies | The chosen destination and the probability of each destination |
| Check a condition | A yes or no question and the evidence | True or false, and the probability of each |
| Apply a policy | The rules and the allowed outcomes | A typed decision based on the text you supply |
| Rate an outcome | Ordered levels, each with clear criteria | The chosen level and the probability of each level |

You supply a context, which is the text to judge, and a schema. The schema must
be flat, which means that it has no nested fields. Each field is an enum or a
boolean. An enum field has a fixed list of string choices, and a boolean field
is true or false.

Each allowed answer has a code that is one token long. The scorer reads the
model's logits for these codes. Logits are the raw scores that the model gives
to each token. The scorer turns the logits into probabilities with the softmax
function. Our Python code then builds the output from these probabilities, so
there is no generated JSON to parse. If a field is an ordered rating scale, your
application can use the probabilities to calculate an expected level.

On a Mac, `ParallelScorer` processes the shared context once and then scores all
the fields in parallel. The CUDA scorer scores each field on its own, with the
full prompt each time. Both scorers return the typed output. They also return
the logits and the probabilities of the candidate answers. Each field is scored
separately, so one field cannot see the answer to another field.

### What you cannot build with the current release

- Nimble accepts only text. You cannot use it to judge other kinds of input,
  e.g., images. This is true even though the base model includes a vision part.
- Nimble only picks from the answers you supply. It cannot write text of its
  own, e.g., an explanation. It also cannot return nested JSON or a piece of
  text taken from the context. An enum field can have 1 to 255 string choices in the latest release,
  and a boolean field has two.
- The probabilities are not a guarantee that an answer is correct. Nimble scales
  them so that they add up to 1 across the answers you supplied. The latest checkpoint uses T=1.0 and has not had a separate temperature fit.
  Earlier releases have their own temperature settings. See [Probability temperature](https://github.com/bespokelabsai/nimble#probability-temperature). A probability
  of 0.9 still does not mean that the answer is right 90% of the time on your
  data. If it is possible that none of your answers fit, add an answer that
  means "no match". Test any probability threshold on your own data before you
  rely on it.
- Each prompt can have at most 8,192 tokens in the latest release. This limit includes the schema and
  the part of the prompt that names the field to score. Nimble rejects longer
  prompts. Fields cannot depend on each other, so your code must check that the
  answers to different fields are consistent.

Nimble’s performance depends on the curated data and domains represented in its training; test it on your own tasks.
But we do see that Nimble is overall better than its base model Qwen3.5-9B in new domains.

## Quickstart

Clone the repository and move into its folder. Run all the commands below from
this folder.

```
git clone https://github.com/bespokelabsai/nimble.git nimble
cd nimble
```

Use Python 3.12. To run the model on a Mac, you need Apple Silicon. Python must
also run directly on macOS so that it can use Metal, which is Apple's interface
to the GPU. To run the model on Linux, you need an NVIDIA GPU that supports
BF16, a 16-bit number format.

Without quantization, the 9B weights alone take about 18 GB. Quantization means
storing the weights with fewer bits to save memory. The model needs more memory
than this while it runs. The merge step below runs on the CPU. It needs extra
RAM, and it needs disk space for both the base weights and the merged weights.
A Mac with 64 GB of memory has more free memory for this than a machine with
24 GB.

### Download the model

Create a Python environment for preparing the model. On Linux, you can also use
this environment to run the model on the GPU. If you do, install a build of
PyTorch with CUDA support that works with your GPU driver.

```
python3.12 -m venv .cache/venvs/nimble
source .cache/venvs/nimble/bin/activate
python -m pip install torch==2.8.0 -r requirements/training.txt
```

The following accepts either a full checkpoint or a PEFT LoRA adapter. For an adapter, it downloads the pinned base and merges the trained weights once. It records the resolved revision and local model path for both platform examples. No TypeSafe or generation API key is needed for local inference.

```
python - <<'PYTHON'
import hashlib
import json
from pathlib import Path

from huggingface_hub import snapshot_download

repo = "bespokelabs/Bespoke-Nimble-9B"  # Or "bespokelabs/Bespoke-Nimble-9B-v2"
snapshot = Path(snapshot_download(repo, cache_dir=".cache/huggingface/hub"))
contract_file = snapshot / "schema_config.json"
contract = json.loads(contract_file.read_text()) if contract_file.exists() else {}
if contract:
    from transformers import AutoTokenizer
    from nimble.training.candidate_schema import validate_contract
    validate_contract(contract, AutoTokenizer.from_pretrained(snapshot))

model_path = snapshot
if (snapshot / "adapter_config.json").exists():
    import torch
    from peft import PeftModel
    from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration

    # The adapter release must include its pinned base and prompt contract.
    base = Qwen3_5ForConditionalGeneration.from_pretrained(
        contract["model"], revision=contract["revision"],
        dtype=torch.bfloat16, device_map="cpu",
    )
    adapter = PeftModel.from_pretrained(base, snapshot)
    merged = adapter.merge_and_unload(safe_merge=True)
    model_path = Path(".cache/models") / ("nimble-9b-" + snapshot.name)
    merged.save_pretrained(model_path)
    (model_path / "schema_config.json").write_text(json.dumps(contract, indent=2))
    AutoTokenizer.from_pretrained(snapshot).save_pretrained(model_path)
    # Preserve adapter identity for automatic temperature selection after merging.
    (model_path / "READY.json").write_text(json.dumps({
        "model": repo, "revision": snapshot.name,
        "adapter_sha256": hashlib.sha256(
            (snapshot / "adapter_model.safetensors").read_bytes()
        ).hexdigest(),
    }, indent=2))

config = {
    "model_path": str(model_path.resolve()),
    "model_id": repo,
    "revision": snapshot.name,
    "max_input_tokens": contract.get("max_length", 2048),
}
Path(".cache/nimble-model.json").write_text(json.dumps(config, indent=2))
print("Ready:", model_path)
PYTHON
```

### Mac with Apple Silicon (MLX)

Use a separate MLX environment after the model preparation step:

```
deactivate
python3.12 -m venv .venv-mlx
source .venv-mlx/bin/activate
python -m pip install -r requirements/mlx.txt
```

In Python, pass the saved model settings to `ParallelScorer` so that it loads
the prepared 9B weights.

```
import json
from pathlib import Path
from nimble.scoring.parallel_scorer import ParallelScorer

config = json.loads(Path(".cache/nimble-model.json").read_text())
scorer = ParallelScorer(**config)
```

Note

If you call `ParallelScorer()` with no arguments, it loads the Qwen3.5-4B
model that we used as a baseline, not Nimble. The MLX runner cannot load a
LoRA adapter folder directly, so use the merged folder that you prepared
above. The MLX runner also does not support quantized weights.

### Linux with an NVIDIA GPU (CUDA)

Activate the environment that you used to prepare 
michael-sumner10
🟧 echo.github ⭐Repo titled 'Data, Model, Recipe for an open Jev': serving reads the prompt once and scores one answer token per question; training fine-tunBespoke Labs (bespokelabsai)β€”β€”
🟠 redditGliner2.5-Decide (Jev style model)
LocalLLaMA
parepeg157
🟧 hnLichen – A local, BOY model system1 (Jev) server with image supportjeff_ciesielski21
🟧 hnJevotron: Multiple Jev integrations from the command linechrismungall121
🟧 hnShow HN: Gutsy, a 0.8B Jev-compatible decision model that runs on your CPUmrkn192
🟠 redditeu/jev - First System One Model Hosted in the EU
LocalLLaMA
juanviera2304
🟠 redditClef: Open Weights decision model by Cloudflare
LocalLLaMA
paf1138403116
🟠 redditPerplexity Decider 27B: Open weights decision model fine tune of Qwen3.8 27B
LocalLLaMA
rm-rf-rm6412
🟠 redditI benchmarked Jev 1.13 against 4 local LLMs on RTX5070 12GB - amazing.
singularity
Storge2130
🟠 redditllama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp
LocalLLaMA
jacek20238231
🟠 redditCalDec v1 - Fully Open Decision Model for Personal Assistants
LocalLLaMA
No_Contract_829629
🟧 hnJev Decision Layer: Save Frontier Tokens on Closed Decisionsvpbhardwaj20
🟧 hnShow HN: Tiny model for fast typed decisions on CPU (Gutsy)mrkn131
🟧 hnGutsy: Tiny model for typed decisions on CPU (Jev style)mrkn122
🟧 hnStartup TypeSafe AI's Jev Model Sparks Copycats, Talk of LLM Alternativesronfriedhaber20
🟧 hnGutsy: Tiny model for typed decisions on CPU (Jev style)mrkn140
🟧 hnShow HN: Gutsy, subsecond decision model inference on CPU (Jev Style), no APImrkn130
🟠 redditI built a lightweight, local Jev-like System One with Ternary-Bonsai-4B β€” and used it as a coding-agent judge
LocalLLaMA
Cultural_Self898000
🟠 redditAplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index
LocalLLaMA
empiriolabsai218
🟧 hnShow HN: TOD,a universal Decision model based on Task Oriented DesignsuriyaG10
🟧 hnJev-Driven SRE Diagnosis: What Worked and What Failedmatt_d5028
🟧 hnShow HN: Tiny model for fast typed decisions on CPU (Gutsy)mrkn120
🟧 hnShow HN: CPU-first model for fast typed decisionsmrkn110
🟧 hnShow HN: Decision Studio – LM Studio for open-source, Jev-style decision modelsdreamlight10
🟧 hnShow HN: CPU First Decision Modelmrkn121
🟠 redditNon-text AI model maker Jev valued at $7.5B just weeks after launch
artificial
lulzxdxdxd10438

Interpretation history

Decision trace