Bespoke Labs released Nimble, an open (Apache-2.0) replication of TypeSafe AI's proprietary Jev β the 'System One' decision model that returns typed choices, booleans, and rubric scores with probabilities instead of generated text. Nimble is a ~165 MiB LoRA adapter on Qwen3.5-9B trained on ~2,676 contrastively curated examples (each flips one fact so the correct answer flips); serving reads the prompt once and scores one answer token per field off the logits, exposed through the same System One request format Jev uses. On its 324-example holdout Bespoke reports 90.1% reference-label agreement vs 93.2% for Jev 1.13.0 and 66.4% for base Qwen3.5-9B, and the release ships the dataset, adapter, prompt builder, and reference inference code. Caveat: the supplied snippets carry only Bespoke's self-reported numbers plus generic Jev explainers (Flowtivity, StackAI, LangChain's harness guide) β no independent verification, no license terms for the code/data, and no third-party adoption of Nimble itself appear in them.
| source | object | author | score | comments |
| π§ hn | Nimble: Data, Model, Recipe for an Open Jev (From Bespoke Labs)Retrieved article excerptOpen article Β· Retrieved 2026-09-30T14:25:27.024691+00:00 # Bespoke Nimble
**Data, Model, Recipe for an open Jev**
[Model](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B) Β· [Updates](https://github.com/bespokelabsai/nimble#updates) Β· [Capabilities](https://github.com/bespokelabsai/nimble#capabilities) Β· [Quickstart](https://github.com/bespokelabsai/nimble#quickstart) Β· [Methodology](https://github.com/bespokelabsai/nimble#methodology) Β· [Documentation and development](https://github.com/bespokelabsai/nimble#documentation-and-development) Β· [Citation](https://github.com/bespokelabsai/nimble#citation)
[Introducing Bespoke Nimble. Serving reads the prompt once and then scores one answer token per question. Data curation changes one fact so that the correct answer flips. Training fine-tunes Qwen3.5-9B with LoRA on the answer tokens only. On 324 held-out examples, Bespoke-Nimble-9B matches 90.1% of the reference labels, compared with 66.4% for its base model and 93.2% for Jev 1.13.0.](https://github.com/bespokelabsai/nimble/blob/main/assets/diagrams/nimble-infographic.svg)
Nimble takes some text and a schema, and makes typed decisions about the text.
The schema is the list of questions to answer. Each question is either a choice
from a list that you give or a true or false question. For each question, Nimble
returns the answer it picked and the probability of each allowed answer.
Nimble makes each decision in one step and does not write out any reasoning
first, so it is fast (blazing fast!). Nimble is
inspired by the System One approach of
[TypeSafe's Jev](https://docs.typesafe.ai/primitives/choice). In this repository,
we share our recipe for training such a model.
Note that we did not distill from Jev. The point of the repository is to show how to curate data, how to train, and to serve such a model, and encourage more research!
You can run [Bespoke-Nimble-9B](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B)
on a Mac with Apple Silicon or on a machine with an NVIDIA GPU.
## Updates
- September 24, 2026: [Bespoke-Nimble-9B](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B) now contains the latest checkpoint, with an 8,192-token context and up to 255 choices per field. Its default is T=1.0; the original release is preserved under the `original-2676` tag, and the separate v2 repository is unchanged. Update this checkout before loading the new release.
- On September 22, 2026, we fitted a temperature for Bespoke-Nimble-9B. With
this temperature, the probabilities better match how often the answers are
right. The model picks the same answers as before. Noul probabilities and
Score values do change, so if you compare them with a threshold, test the
threshold again. See [Probability temperature](https://github.com/bespokelabsai/nimble#probability-temperature) and
[PR #7](https://github.com/bespokelabsai/nimble/pull/7).
- On September 20, 2026, we published the 2,676 training examples and the 324
held-out examples for Bespoke-Nimble-9B. We had left them out of the first
release by mistake. See the [dataset guide](https://github.com/bespokelabsai/nimble/blob/main/docs/DATASET.md) and
[PR #5](https://github.com/bespokelabsai/nimble/pull/5).
- On September 19, 2026, we raised the prompt limit of the hosted API to 8,192
tokens for each question. The model was trained on prompts of up to 2,048
tokens, so shorter prompts are better tested. See the
[SGLang deployment guide](https://github.com/bespokelabsai/nimble/blob/main/docs/MODAL_SERVING.md) and
[PR #4](https://github.com/bespokelabsai/nimble/pull/4).
- On September 18, 2026, [Edgar Dyck](https://github.com/eddited17) added a
public benchmark suite. With it, you can run Bespoke-Nimble-9B and Jev on the
same records from 13 public subsets with human labels. The
[public benchmarks guide](https://github.com/bespokelabsai/nimble/blob/main/docs/PUBLIC_BENCHMARKS.md) has the steps and the
results. See [PR #2](https://github.com/bespokelabsai/nimble/pull/2).
## Capabilities
We built Nimble in one day, so expect some rough edges. What Nimble can do comes
from two sources: the first is the base model, Qwen3.5-9B, the second is our
training data, which we curated for a few specific domains.
### What you can build
| Task | You define | You get back |
| --- | --- | --- |
| Route a request | The destinations and when each one applies | The chosen destination and the probability of each destination |
| Check a condition | A yes or no question and the evidence | True or false, and the probability of each |
| Apply a policy | The rules and the allowed outcomes | A typed decision based on the text you supply |
| Rate an outcome | Ordered levels, each with clear criteria | The chosen level and the probability of each level |
You supply a context, which is the text to judge, and a schema. The schema must
be flat, which means that it has no nested fields. Each field is an enum or a
boolean. An enum field has a fixed list of string choices, and a boolean field
is true or false.
Each allowed answer has a code that is one token long. The scorer reads the
model's logits for these codes. Logits are the raw scores that the model gives
to each token. The scorer turns the logits into probabilities with the softmax
function. Our Python code then builds the output from these probabilities, so
there is no generated JSON to parse. If a field is an ordered rating scale, your
application can use the probabilities to calculate an expected level.
On a Mac, `ParallelScorer` processes the shared context once and then scores all
the fields in parallel. The CUDA scorer scores each field on its own, with the
full prompt each time. Both scorers return the typed output. They also return
the logits and the probabilities of the candidate answers. Each field is scored
separately, so one field cannot see the answer to another field.
### What you cannot build with the current release
- Nimble accepts only text. You cannot use it to judge other kinds of input,
e.g., images. This is true even though the base model includes a vision part.
- Nimble only picks from the answers you supply. It cannot write text of its
own, e.g., an explanation. It also cannot return nested JSON or a piece of
text taken from the context. An enum field can have 1 to 255 string choices in the latest release,
and a boolean field has two.
- The probabilities are not a guarantee that an answer is correct. Nimble scales
them so that they add up to 1 across the answers you supplied. The latest checkpoint uses T=1.0 and has not had a separate temperature fit.
Earlier releases have their own temperature settings. See [Probability temperature](https://github.com/bespokelabsai/nimble#probability-temperature). A probability
of 0.9 still does not mean that the answer is right 90% of the time on your
data. If it is possible that none of your answers fit, add an answer that
means "no match". Test any probability threshold on your own data before you
rely on it.
- Each prompt can have at most 8,192 tokens in the latest release. This limit includes the schema and
the part of the prompt that names the field to score. Nimble rejects longer
prompts. Fields cannot depend on each other, so your code must check that the
answers to different fields are consistent.
Nimbleβs performance depends on the curated data and domains represented in its training; test it on your own tasks.
But we do see that Nimble is overall better than its base model Qwen3.5-9B in new domains.
## Quickstart
Clone the repository and move into its folder. Run all the commands below from
this folder.
```
git clone https://github.com/bespokelabsai/nimble.git nimble
cd nimble
```
Use Python 3.12. To run the model on a Mac, you need Apple Silicon. Python must
also run directly on macOS so that it can use Metal, which is Apple's interface
to the GPU. To run the model on Linux, you need an NVIDIA GPU that supports
BF16, a 16-bit number format.
Without quantization, the 9B weights alone take about 18 GB. Quantization means
storing the weights with fewer bits to save memory. The model needs more memory
than this while it runs. The merge step below runs on the CPU. It needs extra
RAM, and it needs disk space for both the base weights and the merged weights.
A Mac with 64 GB of memory has more free memory for this than a machine with
24 GB.
### Download the model
Create a Python environment for preparing the model. On Linux, you can also use
this environment to run the model on the GPU. If you do, install a build of
PyTorch with CUDA support that works with your GPU driver.
```
python3.12 -m venv .cache/venvs/nimble
source .cache/venvs/nimble/bin/activate
python -m pip install torch==2.8.0 -r requirements/training.txt
```
The following accepts either a full checkpoint or a PEFT LoRA adapter. For an adapter, it downloads the pinned base and merges the trained weights once. It records the resolved revision and local model path for both platform examples. No TypeSafe or generation API key is needed for local inference.
```
python - <<'PYTHON'
import hashlib
import json
from pathlib import Path
from huggingface_hub import snapshot_download
repo = "bespokelabs/Bespoke-Nimble-9B" # Or "bespokelabs/Bespoke-Nimble-9B-v2"
snapshot = Path(snapshot_download(repo, cache_dir=".cache/huggingface/hub"))
contract_file = snapshot / "schema_config.json"
contract = json.loads(contract_file.read_text()) if contract_file.exists() else {}
if contract:
from transformers import AutoTokenizer
from nimble.training.candidate_schema import validate_contract
validate_contract(contract, AutoTokenizer.from_pretrained(snapshot))
model_path = snapshot
if (snapshot / "adapter_config.json").exists():
import torch
from peft import PeftModel
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
# The adapter release must include its pinned base and prompt contract.
base = Qwen3_5ForConditionalGeneration.from_pretrained(
contract["model"], revision=contract["revision"],
dtype=torch.bfloat16, device_map="cpu",
)
adapter = PeftModel.from_pretrained(base, snapshot)
merged = adapter.merge_and_unload(safe_merge=True)
model_path = Path(".cache/models") / ("nimble-9b-" + snapshot.name)
merged.save_pretrained(model_path)
(model_path / "schema_config.json").write_text(json.dumps(contract, indent=2))
AutoTokenizer.from_pretrained(snapshot).save_pretrained(model_path)
# Preserve adapter identity for automatic temperature selection after merging.
(model_path / "READY.json").write_text(json.dumps({
"model": repo, "revision": snapshot.name,
"adapter_sha256": hashlib.sha256(
(snapshot / "adapter_model.safetensors").read_bytes()
).hexdigest(),
}, indent=2))
config = {
"model_path": str(model_path.resolve()),
"model_id": repo,
"revision": snapshot.name,
"max_input_tokens": contract.get("max_length", 2048),
}
Path(".cache/nimble-model.json").write_text(json.dumps(config, indent=2))
print("Ready:", model_path)
PYTHON
```
### Mac with Apple Silicon (MLX)
Use a separate MLX environment after the model preparation step:
```
deactivate
python3.12 -m venv .venv-mlx
source .venv-mlx/bin/activate
python -m pip install -r requirements/mlx.txt
```
In Python, pass the saved model settings to `ParallelScorer` so that it loads
the prepared 9B weights.
```
import json
from pathlib import Path
from nimble.scoring.parallel_scorer import ParallelScorer
config = json.loads(Path(".cache/nimble-model.json").read_text())
scorer = ParallelScorer(**config)
```
Note
If you call `ParallelScorer()` with no arguments, it loads the Qwen3.5-4B
model that we used as a baseline, not Nimble. The MLX runner cannot load a
LoRA adapter folder directly, so use the merged folder that you prepared
above. The MLX runner also does not support quantized weights.
### Linux with an NVIDIA GPU (CUDA)
Activate the environment that you used to prepare | michael-sumner | 1 | 0 |
| π§ echo.github β | Repo titled 'Data, Model, Recipe for an open Jev': serving reads the prompt once and scores one answer token per question; training fine-tun | Bespoke Labs (bespokelabsai) | β | β |
| π reddit | Gliner2.5-Decide (Jev style model) LocalLLaMA | parepeg | 15 | 7 |
| π§ hn | Lichen β A local, BOY model system1 (Jev) server with image support | jeff_ciesielski | 2 | 1 |
| π§ hn | Jevotron: Multiple Jev integrations from the command line | chrismungall | 12 | 1 |
| π§ hn | Show HN: Gutsy, a 0.8B Jev-compatible decision model that runs on your CPU | mrkn1 | 9 | 2 |
| π reddit | eu/jev - First System One Model Hosted in the EU LocalLLaMA | juanviera23 | 0 | 4 |
| π reddit | Clef: Open Weights decision model by Cloudflare LocalLLaMA | paf1138 | 403 | 116 |
| π reddit | Perplexity Decider 27B: Open weights decision model fine tune of Qwen3.8 27B LocalLLaMA | rm-rf-rm | 64 | 12 |
| π reddit | I benchmarked Jev 1.13 against 4 local LLMs on RTX5070 12GB - amazing. singularity | Storge2 | 13 | 0 |
| π reddit | llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson Β· Pull Request #29818 Β· ggml-org/llama.cpp LocalLLaMA | jacek2023 | 82 | 31 |
| π reddit | CalDec v1 - Fully Open Decision Model for Personal Assistants LocalLLaMA | No_Contract_8296 | 2 | 9 |
| π§ hn | Jev Decision Layer: Save Frontier Tokens on Closed Decisions | vpbhardwaj | 2 | 0 |
| π§ hn | Show HN: Tiny model for fast typed decisions on CPU (Gutsy) | mrkn1 | 3 | 1 |
| π§ hn | Gutsy: Tiny model for typed decisions on CPU (Jev style) | mrkn1 | 2 | 2 |
| π§ hn | Startup TypeSafe AI's Jev Model Sparks Copycats, Talk of LLM Alternatives | ronfriedhaber | 2 | 0 |
| π§ hn | Gutsy: Tiny model for typed decisions on CPU (Jev style) | mrkn1 | 4 | 0 |
| π§ hn | Show HN: Gutsy, subsecond decision model inference on CPU (Jev Style), no API | mrkn1 | 3 | 0 |
| π reddit | I built a lightweight, local Jev-like System One with Ternary-Bonsai-4B β and used it as a coding-agent judge LocalLLaMA | Cultural_Self8980 | 0 | 0 |
| π reddit | Aplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index LocalLLaMA | empiriolabsai | 21 | 8 |
| π§ hn | Show HN: TOD,a universal Decision model based on Task Oriented Design | suriyaG | 1 | 0 |
| π§ hn | Jev-Driven SRE Diagnosis: What Worked and What Failed | matt_d | 50 | 28 |
| π§ hn | Show HN: Tiny model for fast typed decisions on CPU (Gutsy) | mrkn1 | 2 | 0 |
| π§ hn | Show HN: CPU-first model for fast typed decisions | mrkn1 | 1 | 0 |
| π§ hn | Show HN: Decision Studio β LM Studio for open-source, Jev-style decision models | dreamlight | 1 | 0 |
| π§ hn | Show HN: CPU First Decision Model | mrkn1 | 2 | 1 |
| π reddit | Non-text AI model maker Jev valued at $7.5B just weeks after launch artificial | lulzxdxdxd | 104 | 38 |
2026-10-10T22:35:44Z
TechCrunch report of Jev's $7.5B valuation confirms the Jev-class category's commercial weight, but does not advance the open-checkpoint adoption question: no production harness has swapped to an open checkpoint, Nimble-specific adoption remains zero, and all live questions (independent Clef/Decider/Aplomb standings, sub-Q8 quantization on official GGUFs, license terms, curation generalization, Jev's calibration edge) are unchanged. Category standardization via llama.cpp mainline /v1/systemone is the strongest adoption signal to date; periphery continues spawn-and-stall.
2026-10-10T21:38:13Z
evidence attached: reddit.post.1x2prfl β TechCrunch reports Jev valued at $7.5B weeks after launch, independent corroboration of Jev-class models' commercial significance
2026-10-10T00:35:18Z
New attachment hn.story.50023015 is a seventh Gutsy re-post (2 pts/1c) β periphery spawn-and-stall pattern continues unchanged. Measured heat shows 0.33 pts/h (61.7 percentile, steady at 226h) but absolute rate remains floor-level; magnitude_valve_eligible reflects the earlier Clef/llama.cpp multi-platform wave already priced in, not new signal. Nimble-specific adoption still zero; llama.cpp mainline /v1/systemone makes swap a config change but none forced. All live questions (independent Clef/Decider/Aplomb standings, sub-Q8 quantization on official GGUFs, license terms, first production-harness swap) unresolved. Case meaning unchanged: open Jev reproducibility established and vendor-standardized; boundary condition recorded (many-fast-low-context gating, not high-context diagnosis).
2026-10-09T19:58:37Z
evidence attached: hn.story.50023015 β shared external link with case evidence
2026-10-08T21:25:31Z
Decision Studio (hn.story.50006805) adds a GUI/workflow tool for open-source Jev-style models β another corroboration of category consolidation, but still no production-harness swap to an open checkpoint. Measured heat at floor (0.17 pts/h, 30th percentile, steady at ~199h). Nimble-specific adoption remains zero; llama.cpp mainline /v1/systemone makes swap a config change but none forced. Live questions (independent Clef/Decider/Aplomb standings, sub-Q8 quantization on official GGUFs, license terms, first production swap) all unresolved. Case meaning unchanged: open Jev reproducibility established and vendor-standardized; boundary condition recorded (many-fast-low-context gating, not high-context diagnosis).
2026-10-08T17:49:20Z
evidence attached: hn.story.50006805 β Show HN for Decision Studio, a GUI/workflow tool for open-source Jev-style decision models, corroborates the trend of open Jev foundations becoming practical agent-harness components
2026-10-08T11:49:40Z
New attach hn.story.50003965 is a seventh Gutsy re-post (1 pt/0c) β periphery spawn-and-stall pattern continues unchanged. SRE deployment report (hn.story.49986765) grew to 50/28 but remains proprietary Jev, not an open checkpoint swap. Measured heat confirms floor-level activity: 0.17 pts/h, steady momentum, 27.5 peer percentile at ~190h age; magnitude-valve eligibility still aggregates the already-priced Clef/vendor/llama.cpp wave. Case meaning unchanged: open Jev reproducibility established and vendor-standardized; Nimble-specific adoption zero; no production-harness swap to any open checkpoint; llama.cpp mainline /v1/systemone makes swap a config change but none forced. Live questions (independent standings, quantization, licenses, first production swap) all unresolved.
2026-10-08T10:42:27Z
evidence attached: hn.story.50003965 β shared external link with case evidence
2026-10-07T10:39:40Z
Both sensor triggers are stale: the 11x velocity multiple is the already-priced SRE deployment report creeping (15/3 β 28/10) against a decayed 0.33 pts/h peer baseline, and the substantive attach is the sixth Gutsy re-post (2 pts/0 c) β periphery spawn-and-stall already on record. Case meaning is unchanged (reproducibility established, vendor wave priced, Nimble-specific adoption still zero, no production-harness swap), so it stays significant/low with material_change false; the 79.6 peer percentile and magnitude-valve flag read as same-age cohort decay plus the already-priced Clef/vendor/llama.cpp spread, not new heat.
2026-10-07T10:25:58Z
evidence attached: hn.story.49990519 β shared external link with case evidence
2026-10-07T05:18:21Z
grounded: converges/high β Converges with his curation-is-the-moat argument (ip:concept.capability-symmetry, ip:source.the-moat-is-the-memory-ebook): a one-day LoRA on 2,676 contrastive e
2026-10-07T05:09:11Z
First applied deployment report in the periphery β 'Jev-Driven SRE Diagnosis: What Worked and What Failed' (15pts/3c) β shifts the periphery's axis from model releases to real-workflow trials and puts a boundary condition on record (System One fits many-fast-low-context decisions; SRE's few-critical-high-context calls are a poor fit), but it runs proprietary Jev rather than an open checkpoint, so it answers none of the awaited signals. Case meaning is otherwise unchanged (category validated and vendor-standardized; Nimble-specific adoption zero; no production swap), so it stays significant/low with material_change false β the 69th peer percentile is same-age cohort decay against a 2.2 pts/h absolute floor, and the magnitude valve still aggregates the already-priced Clef/vendor/llama.cpp wave.
2026-10-07T03:33:14Z
evidence attached: hn.story.49986765 β Applied Jev-class judgment deployment report with honest failure notes is adoption evidence for Jev decision models in real workflows.
2026-10-06T22:52:58Z
Two new periphery entries β Empirio Labs' Aplomb 1 (5.3B, 1M-context, text/image/video/audio in one request, self-run Decision Index 0.2.1 score 44.86, 13pts/7c) and TOD (Gemma-12B scorer + retriever shortlisting, 1pt/0c) β widen the category's capability axes but hit none of the awaited signals; the one genuine nuance is that Aplomb launched against the shared Decision Index scoreboard, the first entrant to treat it as the launch benchmark, a small ecosystem-maturation marker. Case meaning is otherwise unchanged (category validated and vendor-standardized; Nimble-specific adoption zero; no production swap), so it stays significant/low with material_change false β the 64.7 peer percentile is again same-age cohort decay against a 1.5 pts/h absolute floor β and the discrete signals (independent Clef/Decider/Aplomb standings, official-GGUF quantization, license clarity, first production-harness swap) still re-dirty it on arrival.
2026-10-06T20:42:17Z
evidence attached: hn.story.49981275 β Independently open-sourced, self-described Jev-like decision model (Gemma-12B scoring with retriever shortlisting) is direct third-party evidence that open Jev-class decision foundations are spreading.
2026-10-06T20:42:17Z
evidence attached: reddit.post.1wz96qs β New first-party open Jev-class decision model with public Decision Index run β ecosystem data point for whether open Jev foundations become a standard agent-harness component.
2026-10-06T12:13:55Z
Trigger is a qualitatively new periphery addition: an indie builder's Ternary-Bonsai-4B one-pass judge (tree-attention prefix sharing across questions) framed as deployed inside their own coding agent as a judge β the first periphery instance sitting in the target context (agent-harness judging) rather than releasing another standalone model β but it lands with zero traction, no published measurements, and it is the author's own harness, not a production harness swapping decision calls. Case meaning is otherwise unchanged (category validated and vendor-standardized; Nimble-specific adoption still zero), so it stays significant/low with material_change false: low heat reflects zero traction, not a stalled periphery β new authors, base models and use-case framings are still appearing β and the discrete awaited signals (Clef/Decider standings, GGUF quantization, license clarity, first production-harness swap) still re-dirty it on arrival.
2026-10-06T11:33:25Z
evidence attached: reddit.post.1wyzu4k β Independent builder ships an open-weights (Bonsai-4B) Jev-style one-pass judge inside an agent harness β direct third-party adoption evidence for open Jev foundations.
2026-10-05T08:38:08Z
Trigger is Gutsy re-show #5 β same author (mrkn1), 2 pts/0 comments, engagement running 9β3β2β1β2 across five postings with no new measurements β the stalling-copy pattern at its most terminal, so nothing changes the case's meaning. It stays parked at significant/low: the measured 79th-percentile peer read is a same-age cohort decay artifact (peers decay too; the case's hottest object is a Gutsy tick) rather than expansion, and the magnitude-valve flag still aggregates the already-priced Clef/vendor/llama.cpp wave while the only current periphery additions are zero-traction copies; the case awaits its discrete signals (Clef/Decider standings, GGUF quantization, license clarity, first production swap), any of which re-dirties it on arrival.
2026-10-05T08:24:15Z
evidence attached: hn.story.49961971 β shared external link with case evidence
2026-10-04T12:09:11Z
The trigger is Gutsy re-show #4 β same author (mrkn1), no new measurements, engagement decaying 9β3β2β1 pts across the four postings β the stalling-copy pattern at its clearest, so nothing changes the case's meaning. It stays parked at significant/low: the magnitude-valve flag still aggregates the already-priced Clef/vendor/llama.cpp wave while the only current periphery additions are zero-traction copies, and the case awaits its discrete signals (Clef/Decider standings, GGUF quantization, license clarity, first production swap), any of which re-dirties it on arrival.
2026-10-04T11:26:57Z
evidence attached: hn.story.49952885 β shared external link with case evidence
2026-10-04T08:25:40Z
The trigger is the WSJ piece β first mainstream-press coverage of the Jev-alternatives wave, a salience marker that the story escaped the dev bubble, but retrospective coverage of the already-priced September wave: 1 pt/0 comments on HN, no new measurement, adoption, contradiction, or license movement, and none of the discrete signals (Clef/Decider standings, quantization on official GGUFs, license clarity, first production swap) it awaits. The case's meaning is unchanged β category validated and vendor-standardized, Nimble-specific adoption still zero β and the quiet floor (0.17 pts/h vs 154 peak, 35.7th percentile, momentum steady) confirms there is nothing to attend to now; the magnitude-valve flag still aggregates the already-priced Clef/vendor/llama.cpp wave, not ongoing expansion.
2026-10-04T08:23:29Z
evidence attached: hn.story.49951597 β WSJ mainstream coverage of the Jev copycat/LLM-alternatives wave is independent spread evidence for whether Jev-class judgment becomes a standard harness component.
2026-10-03T23:38:31Z
grounded: converges/high β Converges at category level even though Nimble itself still has zero recipe reuse: a one-day LoRA on 2,676 contrastive-curated examples scoring 90.1% of referen
2026-10-03T23:26:46Z
The 10/4 'substantive_evidence' trigger is Gutsy re-show #3 β same author (mrkn1), no new measurements, 2 pts/0 comments β the stalling-copy pattern repeating itself, not new substance, so nothing changes the case's meaning. It stays parked at significant/low: despite the magnitude-valve reading (which still aggregates the already-priced Clef/vendor/llama.cpp wave, not ongoing expansion β the only current periphery additions are zero-traction copies), the live question of which checkpoint captures the open decision-model category and whether any production harness swaps remains genuinely open, awaiting its discrete signals.
2026-10-03T23:25:42Z
evidence attached: hn.story.49948586 β shared external link with case evidence
2026-10-03T12:46:04Z
The three velocity-spike triggers are baseline artifacts β the llama.cpp PR's +5-point drift (77β82) reads as ~20x against a near-zero peer floor while aggregate rate is 0.5 pts/h vs a 147/h peak β and the only new evidence is the same author re-showing Gutsy with no new results, an artifact already catalogued. Nothing moves the case's meaning: it stays parked at significant/low awaiting the discrete signals (Clef/Decider standings, quantization on official GGUFs, license clarity, first production-harness swap) that re-dirty it on arrival.
2026-10-03T12:24:47Z
evidence attached: hn.story.49943105 β shared external link with case evidence
2026-10-02T14:56:13Z
The 'Jev Decision Layer' repo (2 points, 0 comments, unvetted, unclear it even uses an open checkpoint) is the stalling-copy pattern at its thinnest and does not establish the production-harness swap the case waits on β the case's meaning is unchanged. With momentum flipped from steady to cooling, the rate ~8x off peak at ~49h age, and the periphery now producing only zero-traction copies (CalDec, this) rather than new implementations or communities, the episode's attention wave is over: heat drops to low despite the magnitude-valve reading, because that top-decile spread aggregates the already-priced Clef/vendor/llama.cpp wave, not ongoing expansion β while the pending measurement signals (Clef/Decider standings, quantization on official GGUFs, license clarity, first production swap) will arrive as discrete new evidence and re-dirty the case.
2026-10-02T14:26:17Z
evidence attached: hn.story.49933772 β Third-party released Jev decision-layer repo is imitation/adoption evidence for whether Jev-class typed decisions become a standard open agent-harness component.
2026-10-02T13:29:35Z
CalDec v1 is the stalled demo-wave pattern repeating at zero traction β a first-time hobbyist release that itself concedes Jev is the better all-rounder β so the case's meaning doesn't move: a broad but shallow copy ecosystem beneath three real adoption lines (human-labeled verification, two-lab vendor wave, mainline llama.cpp serving), with the live question still which checkpoint wins and whether any production harness swaps. Engagement is frozen across prior evidence (~8x off peak, cooling), but heat holds at medium on the 91st-percentile residual rate and pending Clef/Decider benchmark, quantization, and license signals maturing in days β not because this delta added anything.
2026-10-02T13:24:53Z
evidence attached: reddit.post.1wvtc1j β An independent hobbyist open-weights decision model explicitly benchmarked against Jev β third-party activity in the open-Jev niche that bears on whether open Jev foundations spread.
2026-10-02T11:27:02Z
The llama.cpp /v1/systemone PR with official ggml-org GGUFs for five Jev-class models moves the case's meaning from 'open decision models exist and verify' to 'the default local serving stack now standardizes the model class as a first-class API' β infrastructure-level adoption while Nimble itself still has zero recipe reuse, narrowing the live question to which checkpoint wins and whether any production harness actually swaps. State rises to significant on three independent lines (human-labeled verification, two-lab vendor wave, mainline serving support); heat holds at medium because the numbers run ~8x off peak with steady momentum and the pending Clef/Decider benchmark and quantization signals mature in days, not hours.
2026-10-02T11:23:41Z
evidence attached: reddit.post.1wvqbrz β llama.cpp merging a /v1/systemone API with official ggml-org GGUFs for five Jev-class decision models is direct ecosystem-standardization evidence that open Jev foundations are becoming default harness infrastructure.
2026-10-02T08:26:55Z
The wave has crested into a measurement-waiting phase: engagement is ~7x off peak, no third vendor or new implementation arrived this delta, community fatigue with Jev copies is now voiced, and the only new content is a weak 20-task screenshot benchmark (an unclear community 9B, not clearly Nimble) showing Jev-parity quality at 4x lower latency on an RTX 5070 β the first demand-side datapoint on the local-swap proposition, but too thin to change the case's meaning. Heat drops to medium despite the magnitude-valve spread reading (92nd-percentile rate, 3 platforms) because the periphery stopped expanding within this delta and the pending signals β Clef/Decider benchmark standings, quantization results, license terms β mature in days, not hours.
2026-10-02T08:23:07Z
evidence attached: reddit.post.1wvnrep β Weak but on-point user measurement that a local 9B matches Jev quality at 4x lower latency β the only third-party datapoint in queue on open-Jev reproducibility and Jev's latency moat.
2026-10-02T02:03:43Z
Perplexity's Decider 27B makes the vendor wave two-lab (Cloudflare Clef, then Perplexity), converting last look's 'single strong artifact' caveat into genuine multi-vendor category validation; Nimble's reproducibility thesis is effectively proved out (independent 1.2-point human-labeled gap) while Nimble itself still has zero recipe reuse, so the case's live meaning shifts from 'can open weights reproduce Jev' to 'who captures the open decision-model category' β heat stays high on the expanding vendor periphery (magnitude-valve spread, 98.5th-percentile rate) even though the flagship Clef post is off its peak.
2026-10-02T01:27:07Z
evidence attached: reddit.post.1wvfz9n β Perplexity's open-weights Decider 27B is a second major-lab entry in open decision models, materially contextualizing whether open Jev-class judgment becomes a standard harness component (category validation, not corroboration of Bespoke's claims).
2026-10-01T19:40:34Z
Cloudflare's Clef β open-weights Jev-class post-trains of Qwen3.8-27B and Qwen3.5-9B (clef-flash) β is the case's first influential infrastructure entrant and the closest thing yet to the adoption signal the last look set as its bar: standardization of open Jev-class typed decisions is now being driven at vendor level, not by one-off copies. But the entry is independent of Nimble (no recipe reuse), and the EU 'System One' service landed as a dud (0.07 ratio), so the heat rise is attributable to a single strong artifact rather than broadening periphery; Nimble-specific adoption remains absent and the Jev calibration edge is unclosed.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv4zzi β Independent corroboration: Cloudflare shipping open-weights Jev-class decision models (Qwen3.8-27B and 9B post-trains) with strong community uptake marks Jev-class open foundations going mainstream.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv5vw5 β An independent EU-hosted Jev-class 'System One' decision service signals the open Jev-class ecosystem spreading beyond US vendors, though unvetted (0.1 ratio).
2026-10-01T17:46:52Z
Gutsy repeats the small-open-Jev-model class that Gliner2.5-Decide already opened rather than adding a new artifact class, and lands flat (3/0, no published score) β the periphery's diversification has stalled and mood has turned to copy-fatigue and 'astroturf' hostility, so heat cools to low per the prior look's own condition. The case's meaning is unchanged: proliferation of open Jev-style typed decisions to demo-grade is established (corroborated); what would now move it is a real adoption signal β recipe reuse, a competitive Jev Decision Index entry, or a production swap β not another one-off copy.
2026-10-01T16:32:02Z
evidence attached: hn.story.49923223 β Independent 0.8B open Jev-compatible CPU model corroborates Jev-class structured decisions reproducing on small open weights.
2026-10-01T04:40:46Z
Jevotron β a CLI treating multiple Jev-class providers as interchangeable endpoints β extends the pattern from replication and serving alternatives into commodity plumbing, the shape a standardizing harness component takes; but at 3/0 traction it confirms rather than advances the proliferation priced at the last look, so the assessment itself is unchanged. Platform rates have collapsed (0.17 pts/h vs 22.55 peak, 35th percentile) amid open copy-fatigue; heat holds medium only because the periphery is still producing new artifact classes β count artifacts, not posts β and cools to low if the next look brings no new class or a real adoption signal.
2026-10-01T02:29:13Z
evidence attached: hn.story.49916865 β Independent CLI tooling aggregating multiple Jev-class providers is early ecosystem-adoption evidence for whether Jev-style decisions become a standard agent-harness component.
2026-09-30T23:23:39Z
The case's meaning shifts from 'Bespoke shipped one open Jev replication' to 'open Jev-style typed decisions are proliferating as a harness pattern': within a day, a second independently published open model (Gliner2.5-Decide) reached HF trending and a third-party BYO-model local Jev server (Lichen, with image support) appeared on HN β enough independent implementation lines to pass corroborated. But neither adopts Nimble's recipe itself, each landed with thin traction, and r/localllama shows copy-cat fatigue, so heat holds at medium rather than climbing.
2026-09-30T18:42:40Z
evidence attached: hn.story.49911997 β A third-party local BYO-model Jev server shows the open/local Jev-serving ecosystem forming beyond Nimble and TypeSafe, directly bearing on whether open Jev foundations become a standard harness component.
2026-09-30T18:42:39Z
evidence attached: reddit.post.1wuaq3l β A second independently published open Jev-style model trending on HF corroborates the open-Jev-foundations proliferation the case tracks.
2026-09-30T14:35:34Z
grounded: converges/high β Converges with his curation-is-the-moat position (ip:concept.capability-symmetry, ip:source.the-moat-is-the-memory-ebook): a one-day LoRA on 2,676 contrastive e
2026-09-30T14:27:11Z
case created β First-party release of a complete open Jev replication β data, checkpoint, and recipe, with a live head-to-head benchmark suite against TypeSafe's Jev and a week of dated updates β is a concrete new episode distinct from TypeSafe's own Jev case and the existing local-adapter cases.