2026-10-11 17:10 UTC

TypeSafe AI claims its early-access Jev model delivers frontier-comparable structured decisions with calibrated probabilities at dramatically lower latency and cost than autoregressive LLMs, potentially making real-time software automation cheaper without supporting free-form text generation.

state: resolvedheat: lowuncertainty: lowconvergesscott: highinference-economics frontier-models llm-servingTypeSafe AIDiogo Almeida
Surfaced 2026-09-19T06:22:16Z — priced heat=high at reprice: The expanding implementation periphery and substantial HN/Reddit spread now warrant high attention and accelerating ecosystem status, without establishing Jev’s performance claims. The latest Vampire Survivors integration adds another developer-reported experiment, not a measured result; earlier low heat underpriced the breadth of activity.

What is this?

Jev is the first 'System One' model from TypeSafe AI, a San Francisco startup founded in 2024 and led by ex-OpenAI's Diogo Almeida, which emerged from stealth on 15 September 2026 alongside a $40M seed round led by DCVC at a reported ~$200M valuation. It is a proprietary model that replaces free-form text generation with typed probabilistic answers (choice/score/yes-no plus confidence values) that software can branch on directly; TypeSafe claims frontier-comparable quality at 70–500ms end-to-end and 40–200x faster / 40–400x cheaper than autoregressive LLMs, trained on synthetic data via a proprietary RLCD method, with no architecture, weights, or technical paper published — though the org's own GitHub fork of the LLaDA diffusion-LLM repo is consistent with outside observers' suggestion that it wraps an open-weight base, and its docs confirm text-only native input and shared weights with no per-customer fine-tuning. The supplied web material is entirely launch-side (vendor blog and docs, Business Wire release, Wikipedia stub, Forbes valuation piece) and carries none of the independent record the case has accumulated: third parties confirm bounded accuracy plus a real short-request cost/latency edge, but the load-bearing calibrated-probability claim is refuted by multiple independent ground-truth tests (jevals, a 400-roll fair-die audit, institutional benchmarks), the open-alternative field is saturated at 25+ models, and OpenAI answered the category within two weeks with a Luna-based Decisions API whose pricing, calibration approach, and image-input scope remain unverified.

Why it matters to Scott

The world has independently arrived where Scott's canon already stood and where he is already building: the die-test/jevals/Red Hat refutation of Jev's calibrated confidence lands squarely on his risk-based-triage and deterministic-core doctrine (model confidence is a nudge, never the gate), while OpenAI's Luna-backed Decisions API plus llama.cpp's /v1/systemone and Cloudflare's Clef make the typed-decision front door he already productionised (dev:project.jev — Venture World stage manager, dev-wiki review screener, mail front door, jevkit) an incumbent-native and locally-servable category. That is a dated-receipts publishing opportunity ('the cheap typed front door, weeks before the incumbent shipped the primitive') and a live backend decision: his shortlist (Decisions API vs Jev vs open/local) awaits only the still-unverified Decisions API pricing/calibration specifics, with the jev-gate null and Red Hat result reinforcing his measure-cost-per-accepted-decision doctrine.
dev:project.jevdev:technology.typesafe-jevdev:concept.cheap-model-front-doordev:concept.answers-not-contentip:concept.micro-judgement-patternip:concept.risk-based-triageip:concept.deterministic-coreip:concept.ai-unit-economicsradar:openai-decisions-apiradar:concept.typed-decisionsradar:concept.structured-decisionsradar:concept.model-routingradar:concept.inference-economicsradar:llamacpp-decision-modelsradar:agent-chaperone-jev-tool-screeningradar:typesafe-macos-computer-use
queries asked of Scott's wikis
  • micro-judgement pattern typed decisions routing gates
  • calibrated confidence risk thresholds escalation doctrine
  • decision models vs LLM-as-judge classifier baselines
  • cost per accepted decision inference economics routing
  • logit scoring one-token structured outputs local models
  • jevkit decision backend agent harness integration

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-13 14:00⭐ origin echo-reconstructedAnnounces Jev in early access with parallel typed probabilistic outputs, $0.042 per million input tokens, and workflow evaluations reporting
Diogo Almeida, TypeSafe AI on blog (echo) · attributed from hn.story.49717558
—
09-15 19:25first on hacker news · published · +53.4hJev: New frontier model 40-400x cheaper and 20-200x faster
albelfio
—
09-16 06:00first on r/singularity · published · +64.0hTypeSafe AI releases AI model called Jev. Rather than generating text, it makes decisions. Its hallucination rate is far lower and its outputs are very cheap compared to traditional LLMs.
Profanion
—
09-16 10:39first on r/artificial · published · +68.7hNear Here got early access to TypeSafe Jev, so we tested it for local event validation, tuning each model’s prompt individually. In our tests, Jev delivered up to 5.7× faster responses, 98% lower cost and 12 percentage points higher accuracy - see the results, methodology and limitations
jon_reed
—
09-16 13:34first on r/LocalLLaMA · published · +71.6hLocalJev?
SomewhereAtWork
—
09-17 19:23first on r/ClaudeAI · published · +101.4hI thought I'd found a model 5000x cheaper than Claude for filling web forms. My benchmark was wrong.
imaxalpha
—
09-18 23:54first on r/MachineLearning · published · +129.9hHow is RLCD (jev) RL? [D]
Relative_Wallaby_823
—
09-19 08:25first on r/OpenAI · published · +138.4hI benchmarked Jev aginst gpt-5.6-luna!
LowNefariousness9966
—
09-15 19:25amplified on hacker news 👑hn.story.49717558
albelfio
peak 1908 · 496 comments · 11% of case engagement
09-16 04:47amplified on hacker newshn.story.49722200
handfuloflight
peak 2 · 1 comments · 0% of case engagement
09-16 06:00amplified on r/singularityreddit.post.1whop6b
Profanion
peak 158 · 108 comments · 1% of case engagement
09-16 10:39amplified on r/artificialreddit.post.1whtlzq
jon_reed
peak 2 · 0 comments
09-16 13:34amplified on r/LocalLLaMAreddit.post.1whxf90
SomewhereAtWork
peak 81 · 38 comments · 0% of case engagement
09-16 14:50amplified on r/singularityreddit.post.1whzehf
Oh_boy90
peak 15 · 24 comments · 0% of case engagement
275 more amplifiers in ainews.case_chain
09-15 20:20our radar first saw it · +54.4hdiscovery anchor: hn.story.49717558—
09-19 06:22reached heat=high · +136.4h · via ledger——

Evidence (286) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnJev: New frontier model 40-400x cheaper and 20-200x faster
Retrieved article excerpt

Open article · Retrieved 2026-09-15T21:23:17.811922+00:00

TypeSafe announces System One models and Jev

[Read More](https://typesafe.ai/blog/introducing-system-one-models-and-jev)

[TypeSafe AI](https://typesafe.ai/)

[Manifesto](https://typesafe.ai/manifesto)

[Our Team](https://typesafe.ai/team)

Join Waitlist

TypeSafe announces System One models and Jev

[Read More](https://typesafe.ai/blog/introducing-system-one-models-and-jev)

[TypeSafe AI](https://typesafe.ai/)

[Manifesto](https://typesafe.ai/manifesto)

[Our Team](https://typesafe.ai/team)

[∵ Back](https://typesafe.ai/)

Company News

Sep 14, 2026

# Introducing System One Models and Jev

*Diogo Almeida, founder, TypeSafe*

Models have been superhuman at chat for years, so where is all the automation?

This has been my driving question for the last four years. At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. That work ended up as the research behind ChatGPT.  At the time, I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing.

After two years in stealth, countless technical challenges, and research breakthroughs… I am beyond excited to announce that today, TypeSafe AI is releasing our first **System One Model**: a new class of frontier models built to make fast, structured decisions that software can use directly.

We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).

Our first public model is **Jev**, available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient. While Jev gives up string generation, it’s optimized for structured outputs and *can’t* hallucinate.

Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.

Extraordinary claims require extraordinary evidence so see below for the receipts. 💅

## Frontiers, Old and New

|  | **Existing LLMs** | **System One + Jev** |
| --- | --- | --- |
| Optimized with | Reinforcement Learning with Human Feedback (RLHF) / Reinforcement Learning with Verifiable Rewards (RLVR) | Reinforcement Learning for Calibrated Decisions (RLCD) |
| Optimizes for | Human preference: writeups and chat responses that human raters prefer.    Verifiable rewards: outputs that can be programmatically verified. | Calibrated decisions: answers with epistemically honest probabilities on System One tasks. |
| Inputs | Unstructured data (e.g. text) with an emphasis on **sequential messages**. | Unstructured data (e.g. text) with an emphasis on **structured program state**. |
| Outputs | **Strings / generated text.** Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails. | **Type-safe structured values.** Possible outputs and structure are [defined in advance](https://docs.typesafe.ai/). The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores. |
| Sampling | **Sequential.** Generates one token at a time, each conditioned on the last. | **Parallel.** Generates all outputs in a single query. Incredibly efficient and hardware-aware. |
| Cost | Input tokens: from $0.20 to $10 / MTok.    Output tokens: ~5x more expensive than input tokens. | Input tokens: $0.042 / MTok ($42 per billion tokens).    Output tokens: FREE (too cheap to meter). |
| Speed | **End-to-end response time is** [**3 to 329 seconds**](https://llm-benchmarks.diegoromero.es/)for frontier models.  Fast enough for interfacing with humans, but a big bottleneck when integrated in code. | **End-to-end response time is 70ms-500ms** for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries. |
| Confidence | Even if prompted for a confidence estimate, models tend to be overconfident and inconsistent. If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task. | Always communicates confidence and uncertainty with every output. Calibrated: higher confidence means higher accuracy. More consistent: returns similar answers for similar inputs. |
| Use cases | **Human-in-the-loop tasks (chatbots, copilots, coding agents).** General and powerful, but requires human oversight because their freedom also means they might go off the rails.  **Verifiable problems (math proofs, kernel optimization).** When correctness can be checked cheaply and automatically, LLMs can generate, test, and iterate until they find something that works.    **Demos.** The flexibility of strings allows it to be incredible for quickly making prototypes that only work sometimes. | **AI-Powered Workflows / smart if-statements.** Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems.  **Map-reducing** **over big data.** Turn petabytes of data into features and insights.  **Real-time applications.** 100ms speeds means you can use AI in your applications where UX is critical.**Verify everything.** Score, judge, verify, guardrail, and detect jailbreaks of LLM prompts, reasoning traces, and/or outputs. |

## 

## Evidence / Technical Results

We love skeptics, and are skeptics ourselves.

There are some claims you can easily verify:

- **Speed per call:** We truly are that fast, though our published evals are generally run from our laptops on the West Coast (this is where our service is currently based).
- **Cost per call:** We make our pricing transparent. We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).
- **No type errors**: This would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible.

For our bolder claims, we want to provide as much nuance as we can.

### Side-by-side demonstration

Our side-by-side demo shows a key difference between our models and LLMs: Jev outputs all probabilities in parallel instead of autoregressively generating by token. Strings are extremely powerful and general, but costly. “Giving up” strings actually gives us a lot of superpowers!

[](https://framerusercontent.com/assets/g4FBJpKG573R4zYuoY3HLJIMCu4.mp4)

Nuance:

- For people with early access to TypeSafe, here is the [actual query](https://console.typesafe.ai/playground?share=shr_13a74b495fb786c4bd7964f11597301e7c9).

  - The query is highly simplified and `questions` were chosen to have descriptive, human-readable keys so that the output on the screen is understandable.
  - The `state` is also a short, dense, and detailed paragraph, to emphasize the difference in sampling methodology. The relatively shorter input paints our model in an advantageous light.
- For the keen eyed, for the recorded run, the only disagreement with GPT-5.6 Terra is on “Churn likelihood level”. The actual answer seems genuinely ambiguous to us.
- We used GPT-5.6 Terra with default reasoning for this example, because we’ve found it to be the most comparable at intelligence to Jev on average.
- Fun fact: a similar demo was what convinced us to go all-in in the direction of System One Models!

### Workflow evals

We made a new type of evaluation to measure how well AI works within code. We don’t optimize for a ground truth classification orand allow the harness and model to change (potentially allowing for overfitting via harness engineering). Instead, we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities.

Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

Jev is off the charts – owning the Pareto frontier for almost 2 orders of magnitude. We also compare to models with a generated prompt doing all the logic in their chain-of-thought, but this tends to do significantly worse than using the workflow itself.

Note that the calls here are significantly more complex than the side-by-side demonstration above. That’s because they’re more representative of the types of production workloads needed for true business automation. Below is the simplest of the 4 workflows we’re publishing:

The most reliable real-world workflows tend to have many independent, decomposed questions, with fine-grained behavior that’s dependent on probabilities instead of discrete decisions. The end result is discrete branching, but how we get to a final answer involves a lot of domain-specific engineering that needs to be done highly consistently.

See [our workflow evals site](https://evals.typesafe.ai/) for all the details: examples, disagreements, full queries, and each workflow.

Nuance:

- This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.
- These content of these workflows were not deliberately chosen nor constructed to make our model look good, and are not in our training distribution. However, they were made by individuals on our model capabilities team, so some bias could exist.
- We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models.
- The LLMs use our [System One LLM](https://github.com/typesafe-ai/system-one-adapter-python) wrapper, which constrains LLMs to output structured decisions compatible with our API. We have found this to be the most accurate way to get decisions from LLMs, but this tends to be slower and more expensive than giving decisions without probabilities.

### Hallucination and Type-safety

Hallucination and type-safety are intrinsically related, and we think the latter is table stakes for automation. Having a hallucinated tool call is inconvenient in an agent, but is an absolute deal-breaker if it’s part of a system with latency guarantees or it’s buried several layers deep in a dependency chain. Existing models, *no matter how smart*, still hallucinate and have type errors.

Nuance:

- The numbers for LLMs are from OpenRouter i.e., there almost certainly is bias here: more complex queries might be routed to better models.
- Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.

### Fun Demos

Perhaps the most exciting part of our work is enabling new use cases. We have a lot more to show you, but here are a couple of the team’s favorites:

#### Doom

We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.

[](https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvjpY.mp4)

Nuance:

- The demo is on structured state as a data structure with text, not on images (yet…)
- A non-AI doom bot could play better, but we wanted a bot that was reactive to different representations of game state, and most importantly… following instructions was cool as heck!

#### Wikiracing

The objective of the game is to start on one Wikipedia page and reach a specific ot
albelfio1908496
🟧 echo.blog ⭐Announces Jev in early access with parallel typed probabilistic outputs, $0.042 per million input tokens, and workflow evaluations reportingDiogo Almeida, TypeSafe AI——
🟧 hnGLiClass: Open-Source JEVhandfuloflight21
🟠 redditTypeSafe AI releases AI model called Jev. Rather than generating text, it makes decisions. Its hallucination rate is far lower and its outputs are very cheap compared to traditional LLMs.
singularity
Profanion157108
🟠 redditNear Here got early access to TypeSafe Jev, so we tested it for local event validation, tuning each model’s prompt individually. In our tests, Jev delivered up to 5.7× faster responses, 98% lower cost and 12 percentage points higher accuracy - see the results, methodology and limitations
artificial
jon_reed20
🟠 redditLocalJev?
LocalLLaMA
SomewhereAtWork7738
🟠 redditNew type of LLM released today " it’s optimized for structured outputs and can’t hallucinate"
singularity
Oh_boy901524
🟠 redditQwen3.5 4B + grabbing logits is almost "Jev"? Or even just Qwen Reranker?
LocalLLaMA
theoleecj_n4914
🟧 hnReverse-engineered Jev-like modelrochansinha16023
🟧 hnvLLM: Jev-like mode for the DiffusionGemma modelmmastrac10
🟠 redditI built jev architecture and model one year back for sales
LocalLLaMA
Nandakishor_ml31
🟠 redditI literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper
LocalLLaMA
Nandakishor_ml1544133
🟧 hnJev Ultrafast: A browser agent with a dynamic, indexed action spacerahimnathwani604
🟠 redditI literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper
LocalLLaMA
Nandakishor_ml3351318
🟧 hnOpen-sourced jev architecture last year with model,paper and datasetnandakishor_ml399
🟠 redditIf you have access to Typesafe, try this out
singularity
purealgo43
🟧 hnOpen-jev: One-pass option scoring with Gemma 3 4B, similar to jevsuriyaG41
🟠 redditJev from TypeSafe.ai is getting hyped quite a bit on X. Lots of fun use cases. Not a LLM but a super fast/cheap decision engine with Luna-level intelligence
singularity
manubfr4519
🟧 hnAI Startup launches a faster and cheaper alternative to LLMs for AI automationSarvaturi10
🟧 hnJev for Home AssistantAboveColin10
🟠 redditJev is amazing! I'm letting it play Pokemon Red with a harness being built by Opus 5 in real-time — follow along!
artificial
Boydbme00
🟧 hnJevmlxtrollied20
🟧 hnBenchmark Jev vs. Gemini Flash and Claude Fable on Code Reviewgemanor20
🟧 hnShow HN: Sokit – a LangChain like harness for Jev (or other System 1 models)phantomCupcake20
🟠 redditWhat do you think about Jev and RLCD in general? (Here is my personal take)
artificial
Haghiri7517
🟠 redditI thought I'd found a model 5000x cheaper than Claude for filling web forms. My benchmark was wrong.
ClaudeAI
imaxalpha03
🟧 hnUsing Jev for Claude Code model routinghassleblad2320
🟠 redditjev reproductions tracker. keeping up with jev reproduction efforts
LocalLLaMA
apolinariosteps223
🟠 redditPlaying Doom using Jev by System One, the hottest new thing
singularity
ascii_heart_21
🟧 hnShow HN: Jev routing coding tasks to Grok Build or Codex Astrajoshcsimmons10
🟧 hnIs-odd-jev – check if a number is odd, with a calibrated probabilityalxcrt10
🟧 hnTalk to JEVmkotlikov10
🟧 hnMini-Jev – typesafe's Jev implemented on top of an LLM locallyphyrex30
🟠 redditstill doesn’t get what Jev is…..is it just a more generalised BERT?
LocalLLaMA
AdRepulsive783714547
🟧 hnShow HN: Open-Source Alternative to TypeSafe.aiptitov21
🟧 hnOpen alternative to TypeSafe's Jev, running locally on your own GPUikerM10
🟠 redditMade the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo
LocalLLaMA
Nandakishor_ml771130
🟠 redditNew 'decision' model Jev (developed by co-creator of ChatGPT) is playing Subway Surfers in real-time
singularity
Cagnazzo82547105
🟧 hnProbably – a programming language for LLM workflows, powered by Jevporridgeraisin10
🟠 redditStill on the Jev waitlist? I hosted OpenJev. It's free, go play with it
LocalLLaMA
Every-Comment54737820
🟧 hnOpenJevilreb681284
🟠 redditI use beads with Claude Code, but have found mass labelling of issues inefficient and costly. The new Jev model solves this problem, so I made a tool which uses it to label your beads.
ClaudeAI
bobo-the-merciful14
🟧 hnUlka: Browser agent powered by Jev and fx.sh (experimental)razaan10
🟠 redditI made a Claude Code skill that reaches for Jev before writing another regex, and writes down what happened
ClaudeAI
bobo-the-merciful10
🟧 hnShow HN: Jev vs. GPT-5.6 and Claude Haiku at Pongmatt_oriordan93
🟠 redditJev vs Claude. i built the benchmark so you don't have to
artificial
lutian01
🟠 redditJev is amazing! I'm letting it play Pokemon Red with a harness being built by Opus 5 in real-time — follow along!
singularity
Boydbme5444
🟠 redditJev playing 9 real-time classic games simultaneously with a single API call for $1.80/h
singularity
manubfr12746
🟧 hnNanoJevshock10
🟠 redditDiego Almeida, fondateur de Typesafe AI, présente JEV,
artificial
Winter-Mix-515511
🟧 hnLmjtfy – Ask Jev a yes or no questionrbaudibert61
🟧 hnTypeSafe / Jev latency-focused demos built by Devinrahimnathwani20
🟧 hnShow HN: Jeff – A read-only CLI for semantic code review using Jevimalessandro30
🟧 hnJev as a Primitive Feature of Rubyidz10
🟠 redditRouting between a 3B, a 4B, a 12B and a 26B MoE with Jev > regex
LocalLLaMA
clduab11157
🟧 hnI used Jev to control a swarm of 15 simulated drones in real timekhordoo10
🟠 redditHow is RLCD (jev) RL? [D]
MachineLearning
Relative_Wallaby_8232019
🟧 hnTypeSafe AI's Jev Is Not an LLM – and That May Be the PointBluestein20
🟠 redditDigit-logits-based classifier with llama.cpp
LocalLLaMA
rhinodevil50
🟠 redditUsed Claude to help Jev become a gamer
ClaudeAI
oldmoldycake4612
🟠 redditI benchmarked Jev aginst gpt-5.6-luna!
OpenAI
LowNefariousness996612542
🟧 hnLaya the open source version of Jevnandakishor_ml1326305
🟠 redditHundreds of examples of Jev use cases
singularity
manubfr13434
🟧 hnkev: Jev-like model built on Qwen2.5-0.5Btosh10
🟧 hnA local Jev backed by DiffusionGemmaawei10
🟠 redditVon: Open-source 395M "System One" model
LocalLLaMA
wFXx18277
🟠 redditJev Cuts AI Decision Costs 100x And Vercel, Cloudflare Rushed To Add It
singularity
Prudent-Sorbet-520234680
🟧 hnShow HN: CUA-S1 – A System One Model for Computer Usefrabonacci838
🟧 hnShow HN: VisionLaya: Jev with Vision capabilitiessomeguy10101010
🟠 redditJev vs Luna
OpenAI
aniketmaurya03
🟧 hnIn 2024 I fine-tuned an LLM. Jev could have removed the side questsjc4p21
🟠 redditTypeSafe AI is what the software industry has been waiting for!!!
LocalLLaMA
BoyInDaBox89010
🟧 hnShow HN: S1Code, a decision-first Rust coding agent with Jevmerthdotxyz10
🟧 hnShow HN: Jev to JSON-Schemascosman10
🟧 hnTesting Jev as a validation gate for drug-discovery agentsfred_tandemai30
🟧 hnShow HN: Jev, Fly Me to the Moonrahmanyoo20
🟧 hnShow HN: Jev-align, a CLI to calibrate Jev to your judgementsethkim20
🟠 redditI gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom
LocalLLaMA
shniydder21798
🟠 redditJev demos everywhere — I open-sourced the boring agent-loop recipes for Codex / Claude Code / OpenCode
ClaudeAI
OscarwhDs11
🟠 redditAI Plays Streetfighter 2 In Real Time
OpenAI
Smartaces3938
🟧 hnJev vs. classical ML. Strong on sentiment: Mixed across taskstheanonymousone20
🟧 hnJev is the fastest-adopted model in AI Gateway historyflashbrew21
🟠 reddita local Jev-style decision head onto Qwen 2.5 1.5B
LocalLLaMA
CryOrganic8886162
🟠 redditJev benchmark uses Fable as truth but scores vs Opus
ClaudeAI
Successful-Farm533901
🟧 hnA MySQL plugin that filters rows by meaning (built on TypeSafe Jev)maayanlevy-hn20
🟧 hnDuckDB extension: typed Jev answers as real SQL typesjmrothweiler20
🟠 redditIs Typesafe based/derived from work done by the Laya author?
LocalLLaMA
ECrispy3130
🟠 redditWhat is JEV and what is it used for?
LocalLLaMA
Hot_Example_4456420305
🟠 redditDethrone - a PoC game testing Jev from typesafe.ai, the model that generates decisions instead of text. Evals at the Chronical link at bottom.
artificial
aaddrick10
🟧 hnShow HN: lgtm? – Jev-powered checks that make agents testjennmueng10
🟧 hnAuto approve pull requests with Jevinfiniteregrets20
🟧 hnShow HN: A minimal Pareto-optimal OpenRouter model router for pi, based on Jev7777777phil20
🟠 redditTypesafe's JEV model work as an LLM [P]
MachineLearning
dwarfLevi61
🟧 hnLaya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second)putna15030
🟧 hnJev compiler – Turn the rules your agent keeps ignoring into gates you can testdoronp20
🟧 hnShow HN: Real Jev decisions on a simulated robot fleet – $24.57 per millionchorylee20
🟠 redditlaya.cpp: Optimized laya near-instant decision making
LocalLLaMA
lkarlslund9026
🟠 redditA Jev-style model fine-tuned on Qwen3.5 4B
LocalLLaMA
nato_nob3122
🟧 hnUsing Jev as a teacher to help an SLM write better storiesnutanc10
🟠 redditJev as a classifier is ok but not the part thats important
LocalLLaMA
2BucChuck06
🟧 hnI turned Jev into a (lousy) chatbotkp119716948
🟧 hnSomeone made jev play Atari gamesthomask199530
🟧 hnShow HN: Testing a non-generative decision model on 5,500 CLINC150 inputstgdhtdujeytd30
🟠 redditDIY Jev
LocalLLaMA
Malfeitor12357532
🟧 hnShow HN: jevals – replacing LLM judges with typed Jev decisionsgbayomi281
🟧 hnSub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev0x199720
🟧 hnJev – System-1 Agent Architecture Radar (open-source)noobplus20
🟠 redditconvaiinnovations/laya (multilingual, non-autoregressive System 1 decision model)
LocalLLaMA
Balance-281
🟧 hnJev, Prolog, Pi, and the dream of probabilistic logic programmingschmuhblaster160
🟠 redditI built JevGraph, an open-source pipeline to turn documents into evidence-backed knowledge graphs
LocalLLaMA
richie983007
🟧 hnFind-jevable-code – audit a repo for Jev-replaceable decisionsss_y2n20
🟧 hnKev: Tiny Jev-like family of decision models built on top of Qwen3.5tosh422191
🟧 hnModels Watching Modelstosh10
🟧 hnShow HN: Jeeva – A modular trading engine for mid-frequency trading using Jevmugiwaraa_eth10
🟧 hnShow HN: Grade text from the CLI with custom rulesets and Jevlukstei30
🟧 hnJev and AI SDK Template by Vercel Labsflashbrew10
🟠 redditWhere Jev can take work off Claude and where it cannot, from TypeSafe's own docs
ClaudeAI
prakersh11
🟠 redditA $40M model is being sold on calibrated confidence, and no calibration data has been published
artificial
prakersh911
🟧 hnShow HN: Jev-CLI – CLI wrapper for JEV typesafe AI modeljoshLong14520
🟧 hnI made an LLM using 521 Jev modelsskillseeddev11
🟧 hnWe Tested Jev on 100 Agent Tool Callsarseny_info110
🟧 hnShow HN: OpenDecision – a 400M zero-shot model makes local decisions, plays Doom
Retrieved article excerpt

Open article · Retrieved 2026-09-21T15:25:13.829826+00:00

# OpenDecision

OpenDecision answers typed questions about application state and documents. It runs a local natural language inference model and returns structured values.

## Doom demo

OpenDecision chooses actions for a bot in ViZDoom's Deadly Corridor. This is the Skill 5 recording.

[

Your browser does not support embedded video. [Open the recording](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill5.mp4).
](https://cdn.jsdelivr.net/gh/deepanwadhwa/OpenDecision@main/demos/doom/opendecision-doom-skill5.mp4)

[Watch Skill 1](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill1.mp4) | [Watch Skill 3](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill3.mp4) | [Run the demo](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/README.md)

## Install

With `pip`:

```
pip install OpenDecision
```

With `uv`:

```
uv add OpenDecision
```

[Continue to the get started guide](https://deepanwadhwa.github.io/OpenDecision/quickstart/)

## What it provides

| Type | Use | Result |
| --- | --- | --- |
| `Choice` | Select one option. | Option name and probabilities |
| `Noul` | Test one statement. | A score from 0 to 1 |
| `Score` | Use an ordered scale. | Weighted score and probabilities |
| `Relation` | Compare evidence with a statement and its opposite. | `supports`, `contradicts`, `unknown`, or `conflicted` |
| Document decisions | Ask questions about text or JSON. | Answers and source passages |

[See code examples for each primitive](https://deepanwadhwa.github.io/OpenDecision/primitives/)

## Start here

- [Get started](https://deepanwadhwa.github.io/OpenDecision/quickstart/): install the package, run a Python example, and start the API.
- [Primitives](https://deepanwadhwa.github.io/OpenDecision/primitives/): use `Choice`, `Noul`, `Score`, and `Relation`.
- [Document decisions](https://deepanwadhwa.github.io/OpenDecision/document-decisions/): ask questions about long text or JSON and select a yes/no mode.
- [Evidence and rules](https://deepanwadhwa.github.io/OpenDecision/evidence-and-rules/): rank evidence and combine facts with rules.
- [Examples](https://deepanwadhwa.github.io/OpenDecision/examples/): run the Doom demo and review the insurance and GDPR examples.

## Interfaces

| Interface | Use |
| --- | --- |
| Python | Call OpenDecision in the same process as the application. |
| `POST /v1/systemone` | Send state and typed questions to the API. |
| `POST /v1/documents/decide` | Send a document and typed questions to the API. |
| TypeSafe-compatible endpoint | Use a compatible TypeSafe SDK client with a local server. |

## Basic example

```
from opendecision import OpenDecisionEngine

engine = OpenDecisionEngine()

result = engine.choice(
    state="The customer was charged twice.",
    instructions="Which team should handle this request?",
    criteria={
        "billing": "Payments, invoices, refunds, and duplicate charges",
        "technical": "Software bugs",
        "sales": "Pricing and purchases",
    },
)

print(result["choice"])
# billing
```
dwa359260
🟧 hnShow HN: Run Jev-style models locally on Mac with 0.74 GB RAM
Retrieved article excerpt

Open article · Retrieved 2026-09-21T15:25:28.263414+00:00

# Laya MPS

**Run Jev-style typed decisions locally on your Mac with low RAM usage and fast responses.**

Laya delivers typed decisions with ~32 ms median latency using ~2.1 GiB RAM on
M5 Pro, with a slower ~0.74 GiB mode for lower memory use.

[Laya MPS Pong demo with live response latency](https://github.com/afshinm/laya-mps/blob/main/.github/assets/pong-demo.gif)

MPS stands for [Metal Performance Shaders](https://developer.apple.com/metal/pytorch/),
which PyTorch uses to run Laya on your Mac's GPU.

[Laya typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions)
chooses options, scores inputs, and estimates whether statements are true. The
English model specializes in customer service, invoices, security incidents,
and agent traces. It is not a general-purpose language model.
[Model comparison and benchmarks](https://github.com/afshinm/laya-mps/blob/main/benchmarks/README.md).

## Requirements

- Apple Silicon Mac (M1 or newer), macOS 14+.
- Git and [uv](https://docs.astral.sh/uv/getting-started/installation/). uv installs Python 3.12 if needed.
- Allow 4 GB of free disk space for the runtime, download cache, and ~843 MB model.

## Get started

Clone the repository and start the server:

```
git clone https://github.com/afshinm/laya-mps.git
cd laya-mps
./scripts/serve.sh
```

The first run installs dependencies and downloads the model. Open
**[the demo](http://127.0.0.1:8000/demo/)** and click **Play** or **Run Benchmark**.
The benchmark runs for 60 seconds and plots response latency.

The server runs on `127.0.0.1:8000`. Later runs use local files and inference
works offline. Stop it with **Ctrl+C**. To update a clone, stop the server,
run `git pull`, then run the script again.

## Memory settings

All settings run the same complete model in FP32. Measured on an **M5 Pro,
24 GiB RAM**, macOS 26.4:

| Setting | Peak process RAM | Median decision latency |
| --- | --- | --- |
| `minimal` | **0.74 GiB** | 171 ms |
| `reduced` (default) | 2.11 GiB | **32 ms** |
| `full` | 2.30 GiB | 32 ms |

RAM is the peak across the benchmark suite, excluding macOS, the browser, and
other apps. Latency uses fixed Pong inputs and excludes HTTP. All settings
matched exactly on **260/260 decisions**, including probabilities.
[Full measurements](https://github.com/afshinm/laya-mps/blob/main/benchmarks/README.md#memory-and-latency).

The default keeps transformer layers in RAM and reads embeddings from disk.
`minimal` also streams layers from disk; `full` keeps all weights in RAM.
To change the setting, stop the server and restart:

```
./scripts/serve.sh --memory minimal
```

How memory savings work

Resident weights load directly into their final FP32 allocation, avoiding a full
host-model copy. The checkpoint stores 16-bit values; expanding them to FP32
preserves their values. Computation stays in FP32 in every mode.

Disk embedding lookups fetch only the needed token rows, combine adjacent reads,
and restore the original token order. A 1 MiB row cache keeps the 196.75 MiB FP32
embedding table off the GPU in `minimal` and `reduced`.

`minimal` executes all 28 encoder layers and two decision-head layers through
shared buffers: 24.03 MiB on the host and 48.05 MiB on the GPU. Each forward
requests about 702.66 MiB of layer bytes. GPU work finishes before buffer reuse.
Multiple questions can require multiple forwards, which adds disk-read latency.
`F_NOCACHE` is requested on macOS, but OS caches may still serve reads; logical
read volume is not physical SSD traffic.

The runtime defaults `PYTORCH_MPS_LOW_WATERMARK_RATIO` to `0.00001` before GPU
allocation to encourage smaller Metal heaps and earlier reclamation. Explicit
environment values take precedence. This is a soft watermark, not a RAM limit;
the hard high watermark stays at its runtime default. When using Python directly,
create the engine before other MPS allocations. Effective settings appear in
`diagnostics=true` responses.

Implementation references: [Laya source](https://github.com/NandhaKishorM/laya/tree/6a5819129eb220570792e417e49723d697efd76f),
[MPS allocator](https://github.com/pytorch/pytorch/blob/08187d9e0fba026dc8217405802ab5381dc88d90/aten/src/ATen/mps/MPSAllocator.h),
[PyTorch settings](https://docs.pytorch.org/docs/2.14/mps_environment_variables.html),
[Safetensors format](https://github.com/huggingface/safetensors#format).

## Use the API

With the server running, choose a team for a support ticket:

```
curl --fail-with-body -sS http://127.0.0.1:8000/v1/decisions \
  -H 'Content-Type: application/json' \
  --data-raw '{
    "state": "The export page returns HTTP 500. Our team cannot finish its work.",
    "questions": {
      "owner": {
        "type": "choice",
        "instructions": "Which team should investigate this report?",
        "criteria": {
          "billing": "Payments and invoices",
          "engineering": "Software failures",
          "unknown": "Insufficient information"
        }
      }
    }
  }'
```

Example response:

```
{
  "model": "convaiinnovations/laya-typed-decisions",
  "answers": {
    "owner": {
      "type": "choice",
      "choice": "engineering",
      "probabilities": {
        "billing": 0.06904838234186172,
        "engineering": 0.6391704082489014,
        "unknown": 0.29178112745285034
      },
      "confidence": 0.24445859127951464
    }
  }
}
```

Check whether work is blocked and score the impact in one request:

```
curl --fail-with-body -sS http://127.0.0.1:8000/v1/decisions \
  -H 'Content-Type: application/json' \
  --data-raw '{
    "state": "The export page returns HTTP 500. Our team cannot finish its work.",
    "questions": {
      "blocked": {
        "type": "noul",
        "instructions": "Does the report say that work cannot proceed?"
      },
      "severity": {
        "type": "score",
        "instructions": "Rate the operational impact described in the report.",
        "levels": [
          "Cosmetic issue with no work affected",
          "Some inconvenience but work can proceed",
          "Work is blocked for the whole team"
        ]
      }
    }
  }'
```

Example response:

```
{
  "model": "convaiinnovations/laya-typed-decisions",
  "answers": {
    "blocked": {
      "type": "noul",
      "noul": 0.7066293954849243
    },
    "severity": {
      "type": "score",
      "score": 1.786898910999298,
      "probabilities": [
        0.019407516345381737,
        0.17428618669509888,
        0.8063063621520996
      ],
      "legend": {
        "0": "Cosmetic issue with no work affected",
        "1": "Some inconvenience but work can proceed",
        "2": "Work is blocked for the whole team"
      },
      "confidence": 0.4951949684505006
    }
  }
}
```

Responses contain `model` and typed `answers`. `choice` selects a label, `noul`
is the probability of true (0–1), and `score` is a probability-weighted level
number (0–2 in this example). Confidence describes how concentrated the
probabilities are; it does not establish correctness. Evaluate the model on
your own inputs before relying on it.

Timing and diagnostics

Add `?metrics=true` to include engine evaluation time and current server-process
RAM. The demo measures full HTTP response time separately. Memory is physical
footprint on macOS, or RSS on CPU platforms without that measurement; it is
not peak RAM or whole-machine RAM.

```
curl --fail-with-body -sS 'http://127.0.0.1:8000/v1/decisions?metrics=true' \
  -H 'Content-Type: application/json' \
  --data-raw '{
    "state": "The export page crashes. Our team cannot finish its work.",
    "questions": {
      "blocked": {
        "type": "noul",
        "instructions": "Does the report say that work cannot proceed?"
      }
    }
  }'
```

Example response (timing and memory vary):

```
{
  "model": "convaiinnovations/laya-typed-decisions",
  "answers": {
    "blocked": {
      "type": "noul",
      "noul": 0.6566364169120789
    }
  },
  "metrics": {
    "request_ms": 20.240125013515353,
    "memory_bytes": 2068972792
  }
}
```

Add `?diagnostics=true` for token counts, temperatures, logits, native
probabilities, and auxiliary action probability, plus runtime versions,
timing, memory snapshots, and storage I/O. Both query options can be combined.
Native Noul probabilities are ordered `[false, true]`; the public `noul` value
is the probability of true. The auxiliary action probability is a separate head.


API reference and limits

Base URL: `http://127.0.0.1:8000`. No API key is needed. The server accepts local
connections; cross-origin browser access is not enabled.

| Endpoint | Returns |
| --- | --- |
| `GET /health` | `{"status":"ready","busy":false}` |
| `GET /v1/config` | Model, revision, memory setting, device, precision, and limits |
| `POST /v1/decisions` | Model identifier and typed answers |
| `GET /docs` | Interactive API reference; UI assets load from a CDN |
| `GET /openapi.json` | API schema, available offline |
| `GET /demo/` | Pong and the latency benchmark; no CDN assets |

Send `Content-Type: application/json`. `state` accepts finite JSON and
`questions` contains 1–16 named questions. Optional top-level `instructions`
apply to every question. Questions run independently.

| Type | Input | Answer |
| --- | --- | --- |
| `choice` | `criteria`: 2–26 labels with descriptions | Label, probabilities, confidence |
| `score` | `levels`: 2–10 descriptions, low to high | Weighted zero-based score, probabilities, legend, confidence |
| `noul` | Optional `criteria` with `false` and/or `true` descriptions | Probability of true |

Instructions and descriptions are strings. Score probabilities follow level
order. The default model allows 1,024 formatted tokens per question; the optional
`english` checkpoint allows 512. Formatting includes the state, instructions,
and options. Inputs that would be shortened are rejected.

| Status | Meaning |
| --- | --- |
| `400` | Invalid host, encoding, or unparseable body |
| `413` | Body exceeds the default 1 MiB limit |
| `422` | Invalid fields, options, JSON, or context overflow |
| `503` | Inference is busy or model execution failed |

One inference runs at a time. A busy response includes `Retry-After: 1`; health
and configuration remain available. Runtime settings are chosen at startup.
The `X-Laya-MPS-Config` header identifies the active configuration on config and
decision responses, allowing the demo to detect changes during a run.

The answer types follow [Jev's typed decisions](https://docs.typesafe.ai/introduction/quickstart).
This API has its own request limits, Score `levels` field, and probability-list
format; it is not a drop-in Jev endpoint.

## CLI and setup

Save the JSON payload from a curl example as `request.json` to evaluate it
without starting a server:

```
uv run --locked laya-mps decide request.json --download
uv run --locked laya-mps decide request.json --metrics
```

Pass server options to `scripts/serve.sh`. Use `--port 8001` if port 8000 is busy,
`--checkpoint english` to evaluate the original English model, or
`--model-dir PATH` to use another model directory. `--device cpu` is an explicit fallback;
the published performance numbers use the Mac GPU. The default
`--question-batch-size 1` limits peak RAM. List all options with
`uv run --locked laya-mps serve --help`.

Setup checks and offline startup

```
uv run --locked laya-mps doctor
uv run --locked laya-mps setup
UV_OFFLINE=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 ./scripts/serve.sh
```

`doctor` checks GPU availability and model readiness. `setup` downloads the
pinned model revision. Models live in `.models/` unless `--model-dir` is set;
the helper script always runs from the repository root. `serve --download`
repairs incomplete installations and reuses complete ones. The setup marker is
written only after all required files are present. There is no automatic model
or device fallback.

Use a normal terminal if a restricted environment cannot access Metal. Setup
requires internet access and Git; t
afshinmeh30
🟧 hnJevEmon: Typed Decisions over GBA RAM to Walk Pokémon FireRed
Retrieved article excerpt

Open article · Retrieved 2026-09-21T15:25:29.148006+00:00

# JevEmon

Jev walks a real Pokémon FireRed ROM.

**Bring your own ROM, we don't ship it.**

Not a screenshot agent. Not a bot that mashes A. Each leg of the walk, the code reads the overworld out of RAM, works out every place you could actually go from here — a door, a path to the next route, a Pokémon Center if your party needs one — and hands that list to [Jev](https://docs.typesafe.ai/models.md). Jev picks a destination; the code paths there and presses the buttons. If a wild Pokémon interrupts the walk, Jev fights it out with the same kind of typed decision, then the journey re-plans from wherever the encounter left you.

More on how this is wired: [ARCHITECTURE.md](https://github.com/daniel4x/JevEmon/blob/main/ARCHITECTURE.md)

[Jev walking from the Player's House through Pallet Town to Viridian City](https://github.com/daniel4x/JevEmon/blob/main/docs/journey.gif)

## Status

Verified milestones so far: Jev can leave the Player's House, cross Pallet Town, deliver itself through Route 1 (fighting and winning any wild encounters along the way), and reach Viridian City. Further legs of the journey (Oak's Parcel, the first Gym) aren't built yet — the walk currently ends the run once it reaches Viridian City.

## Try it

Grab a [TypeSafe API key](https://typesafe.ai) and put it in `.env`.

```
cp .env.example .env
```

Drop `Pokemon - Fire Red Version (U) (V1.1).gba` next to this README. Requires a Mac with Apple Silicon.

```
brew install mgba ffmpeg uv
uv run python scripts/setup.py
uv run python -m jevemon --prepare
uv run python -m jevemon
```

That builds the checkpoint, starts the local server, and opens <http://127.0.0.1:8765> for you. Hit Go live and sit with it. Edit the lineup in the page. Speed it up. Scrub the VOD after. Everything runs through `uv` — no shell wrapper scripts, no separate install step.

One walk, no browser:

```
uv run python -m jevemon --run --speed 4
```

Other useful commands:

```
uv sync                                    # just the Python environment
uv run python -m jevemon --check-journey   # offline, no API calls
uv run python -m jevemon --port 9000       # serve on a different port
```

## What Jev sees

Not the screen.

A JSON snapshot of the overworld: where you are, your party's HP, and every destination the pathfinder already confirmed is reachable from here — each with how many times you've already visited it. During a wild encounter, the snapshot switches to the same battle state as any other fight: HP, types, moves, stats, abilities, PP, and only the legal moves and switches for that turn.

Jev answers with one choice and a probability for every option offered. If the answer is illegal, the run stops. There is no backup brain.

On the ROM, a switch is the Pokémon's personality value, not "slot 3." The party menu moves around. Using the old slot would send out the wrong one. That note lives in [AGENTS.md](https://github.com/daniel4x/JevEmon/blob/main/AGENTS.md).

Edit the starting lineup in `config.json` or the UI.

## This is a demo

The interesting part isn't the exact team or the exact route. It's that a typed choice over honest game state — no screenshots, no free-text prompting — is already enough to walk an overworld and handle whatever interrupts it.

Fork it. Push the journey further. Teach it a destination beyond Viridian City, or a whole different region.

Third-party dependencies and their licenses are in [NOTICE.md](https://github.com/daniel4x/JevEmon/blob/main/NOTICE.md).

## License

Copyright © 2026 Daniel Alfasi. Licensed under the GNU General Public License v3.0 — see [LICENSE](https://github.com/daniel4x/JevEmon/blob/main/LICENSE).
alfasiii10
🟠 redditA weekend with Jev made my coding agents up to 31% faster
ClaudeAI
bartlomein010
🟧 hnUsing Jev as an LLM Linter for OMP (Pi) Agentgoulinkh10
🟧 hnUsing TypeSafe AI in Bug Bountylampysecurity10
🟠 redditOn a small pilot, Gemma 4 E2B's verbalized confidence put 99.6% of its judgments on 0.00, 1.00 or 0.50, and a separate judge (Jev & Laya-type arch) ranked the same passages better, AUROC 0.926 vs 0.804, including the ones Gemma called certain
LocalLLaMA
clduab11013
🟠 redditJev vs Laya head to head benchmark [D]
MachineLearning
bobo-the-merciful10
🟠 redditLLMs can already mostly do what Jev does if you limit them to 1 token of output
singularity
arkuto014
🟠 redditI went through 600+ Jev builds. the interesting part isn't the flashy demos
artificial
Sarthak999gupta04
🟠 redditExploring small LLMs as classifiers to rival Jev (sharing what I've found)
LocalLLaMA
OneFanFare87
🟠 redditJev's calibration was measured. The LLMs won [D]
MachineLearning
frappuccinoCoin01
🟠 redditUsing Jev to automatically moderate social media posts according to your site's rules.
OpenAI
RealDannyhvv010
🟧 hnJev: System One Models for Prod, Not God – With Diogo Almeida, CEO, TypeSafe AIswyx41
🟠 redditDeepSeek 4.1 Flash (as a System One model) vs Jev
LocalLLaMA
frappuccinoCoin08
🟧 hnShow HN: Jevopt: Making intelligent compiler optimisation decisions with Jevramneet_singh10
🟠 redditKev: tiny Jev-like decision models (0.8B/4B/9B) on Qwen3.5 you can train and run locally - the 9B fits a 32GB Mac
LocalLLaMA
khiladi17292210
🟠 redditjevals: locally runnable evals for agents using Jev-style decisions
LocalLLaMA
byebaybay10
🟠 redditJev is very good. Your own data is better.
LocalLLaMA
TrifleHopeful5418014
🟧 hnLaya vs Jev head-to-head on identical inputsbeckford10
🟠 redditJev to autoselect Claude model - co-creator of chatgpt's project
ClaudeAI
fsharpman912
🟧 hnJev introduces a new shape of LLMbenwerd4213
🟧 hnjev-router: route to the cheapest model in claude code for your tasksaikatsg20
🟧 hnTypeSafe AI Jev vs. GPT-6 Astraflashbrew40
🟧 hnAutoresearch on Jev: improving a Jev harness to beat Deekseekscosman20
🟧 hnWhere Jev worked for us, and where it didn'taozisik20
🟧 hnShow HN: Blink – A high-performance Jev like decision model for C and WASMmarcobambini10
🟧 hnShow HN: Gemma 3 4B as a typed decision function in Rust (47 ms/decision)zozo123-IB-IL210
🟠 redditQevi-2B: A Jev-style finetuned model for image classification
LocalLLaMA
Taronyuuu00
🟧 hnOpenAI is about to eat Jev's lunch – Arcturus LabsJohnBerryman299212
🟠 redditI reverse-engineered Jev and rebuilt it as open weights. It beats mine on all 7 benchmarks; mine beats it on my own task with 395 labels.
LocalLLaMA
s1lv3rj1nx015
🟧 hnI deleted slow, expensive LLM turns in coding agents with Jevrespectattentio10
🟧 hnI built a linter/skill that ensures requests to jev are structured correctlysuraj_phanindra10
🟧 hnWe put Jev in production against a cross-encoder. Here are the numbersdennispi20
🟧 hnGo-System-One: Jev-Like Results in Pure Go and SIMD/PTXrcarmo11
🟧 hnShow HN: Ego-jev – 0.4s typed decisions for browser agentsZephyrDeng10
🟧 hnCalibrating Jev as a Code ReviewerLectem10
🟧 hnJevBench, a reproducible benchmark for typed decision modelsflorianstandhar11632
🟠 redditJevBench
LocalLLaMA
openSourcerer900003
🟠 redditDoes Jev reveal hidden sexist and racist tendencies in AI?
artificial
No-Wishbone2391020
🟧 hnFaster and local Jev like model for Macmkagenius10
🟧 hnTinyJev -Tiny Jev-style decision model that runs offlineaglaweankit20
🟠 redditIs jev, text embeddings but without chunks? (results included)
LocalLLaMA
No_Afternoon_426038
🟧 hnI built the Jev architecture one year ago and open-sourced ithtk191
🟠 redditI tested Semif (openjev) vs Von vs Jev
LocalLLaMA
KingPinX19
🟠 redditUsing Jev to evaluate LLMs, RAG, and agents - not just give them a score
LocalLLaMA
Charming_Group_295000
🟠 redditstuntd: a local Jev-compatible server on Laya that learns from your own traffic (no API key needed)
LocalLLaMA
Inevitable-Log54142410
🟧 hnJevify skill – Gets your existing agents running on Jevsoupz01120
🟧 hnJev in 25 Lines of Pythonbashbjorn576186
🟠 redditJev in 25 lines of Python
LocalLLaMA
johnnyApplePRNG269101
🟠 redditA proxy that watches the yes/no and pick-one decisions your app asks an LLM for, then trains a local model to make them for free
artificial
Inevitable-Log541414
🟠 redditKev - an open-source System One decision engine that can compete with Jev, using a local 9B model
ClaudeAI
Miserable_Extent884513
🟠 redditI turned Qwen3.8-27B Q2_64 + llama.cpp into a fully TypeSafe AI-compatible Jev-like system. OpenAI API still intact! World’s first Vision-enabled Jev-like model! <10 GB VRAM, 170 ms on an RTX 3090 and ~140 tok/s in chat. 76% vs. 88% Jev-1.13 Acc. on a diverse 22,000-request typed-decision benchmark
LocalLLaMA
kyr0x0029
🟠 redditMods: can we do something about half the forum getting filled with these advertising posts for Jev?
LocalLLaMA
Acrobatic_Stress13881182240
🟧 hnBeating Jev's accuracy, speed, and cost with open modelsrob31320
🟧 hnLaya MPs Sourcepiqufoh10
🟧 hnJev does not play dicesomeguy10101010
🟧 hnJev Can't Be Calibratedalexmolas914
🟧 hnI Couldn't Build Jev at OpenAI – Diogo Almeida, TypeSafe Co-Founder and CEO [video]ABS40
🟧 hnShow HN: Open Code for Jevphegler10
🟠 redditI made a Go adapter that gives local llama.cpp models a Jev-like decision API
LocalLLaMA
wenyani019
🟠 redditThis time I tested reflex vs SemIf vs Laya vs Von vs jev
LocalLLaMA
KingPinX23
🟠 redditJev isn't new tech. Its marketing targets people who think AI started with LLMs.
LocalLLaMA
tiensss773288
🟠 redditjevals: locally runnable evals for agents using Jev-style decisions
LocalLLaMA
byebaybay04
🟧 hnJev vs. LLMs on 770 "Am I the Asshole?" postsdchristopoulos162
🟧 hnJev is 13.6x faster, 2.7x cheaper than GPT Luna 6sjmaplesec22
🟠 redditJEV broke down 724 live ads from 37 brands in 40 seconds for $0.09 of tokens
artificial
sibraan_010
🟧 hnJev deserves hype but not the type its gettingyididev11
🟧 hnI tried using Jev for judgment based guardrailsdeepanshsaxena20
🟠 redditI made deepmoney a while back. My new side project: using the stock market as the label for a Jev-style news screener on Qwen3.8-27B
LocalLLaMA
Fun_Water223002
🟠 redditJEV almost dead: CLM vs JEV
LocalLLaMA
R_Duncan422180
🟧 hnContrastive Language Model (CLM): An Ultra-Fast System One Modelpiyushsthr20
🟠 redditJev using my computer to make a linkedin post
singularity
bGivenb1523
🟧 hnWhat Is RLCD? The Secret Behind Jevtnspacetime638
🟧 hnJev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die rollkantahayashi49
🟠 redditI built an open-weight alternative to Jev / TypeSafe - introducing OpenJudgement-4B (early preview)
LocalLLaMA
bakatristan1413
🟧 hnJev Against a Cross-Encoderfelineflock30
🟧 hnLocal JEV, fine tuned in < 30 mins on Mac Airshmc10
🟧 hnShow HN: Cbjev – typed decisions about text from one encoder passtomek766711
🟧 hnJEV-Star: Low-Cost StarCraft II Control with Language-Model Planningtndl10
🟠 redditOpus 5.5 co-designed and ran a benchmark and evaluation of jev-1.13
ClaudeAI
giulioc8412
🟧 hnShow HN: Jev-pilot – Jev picks Claude Code's effort, model and skill per promptakramovic20
🟧 hnValen: A multimodal decision model inspired by Jevtaylorfinley10
🟧 hnShow HN: Jev-browse – browser sub-tasks for coding agents at ~1/3 the costdanielnc10
🟧 hnJev and System One Models: Calibration Beats Accuracypansuriyakartik1311
🟧 hnJev Based Code Reviewnamanbhulawat4655
🟧 hnPerch: Semantic Code Linting with Jevhandfuloflight10
🟧 hnShow HN: Jevper – an LLM API client constrained to the Jev wire formatczl_my10
🟧 hnShow HN: Knowledge Signal – A JEV-powered rubric assessment tool for study notespaperplaneflyr20
🟠 redditJev didn't beat Claude, it just made me realize how much I overuse Claude
ClaudeAI
Intrepid_Truth889814
🟧 hnShow HN: JevPertus – Jev-style option scoring on Apertusjaluus20
🟧 hnI made a linter using Jev for things a regular linter can't catchczxtm20
🟧 hnA Jevlike using an escalation model to route between a classifier and an LLMyogthos20
🟧 hnReAnchor, a Jev Optimizer in DSPydbreunig10
🟧 hnBenchmarking Jev, Laya, and five open models under three stresseslovegreenlife30
🟧 hnShow HN: Jev Plays Pokémon Redpancomplex272121
🟠 redditJev vs. Kev: open-source Jev alternative tested side by side
LocalLLaMA
facethef11842
🟧 hnJev vs. Kev: open-source Jev alternative tested side by sidefelix089122
🟧 hnOllaya – Ollama for open-source, Jev-style decision modelsArdakilic578139
🟧 hnContrastive Language Models: A Fast, Generalizable System One Modelyarapavan10
🟠 redditJev Playing Pokemon Red Live (Open Source)
singularity
supportingthedogs00
🟠 redditKev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end
LocalLLaMA
Top-Evidence174347
🟠 redditMica v0.1 4B got an iron pickaxe in real Minecraft without generating a single token
LocalLLaMA
Top-Evidence17427443
🟠 redditMica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time
LocalLLaMA
Top-Evidence174246
🟧 hnJev4jjsumrall11
🟠 redditCLM-v0.1-8B ported to MLX — frozen Qwen3-8B encoder for instant on-device decisions, 99% top-1 agreement with the original vLLM server
LocalLLaMA
WebAssemblyMan01
🟧 hnTyped-lm: a Rust jev open source alternativeandrelgcclaudin31
🟠 redditAn open-source alternative to Jev
artificial
IceBergRock546
🟠 redditIs Jev worth the hype?
singularity
beasthunterr693857
🟧 hnShow HN: Ephemeral runner for JEV-style modelszatsepin10
🟠 redditJev already has an open-weight competitor - Deem 9b
singularity
Pokenhagen8320
🟧 hnTurning GLM-5.3-Flash into a Jev-like decision modelflxflx12958
🟧 hnVon, an Open-Source Jev Alternativek__20
🟧 hnShow HN: A local alternative to Jev – 94% on Banking77nico10
🟧 hnCredence – Jev-style typed decisions from a local GGUF modelnaughtiusmax10
🟠 redditGLM-5.3-Flash works as a Jev-like decision model with the same accuracy and speed
singularity
yogthos8616
🟠 redditSupersonicLabs/Julia-1 · Hugging Face
LocalLLaMA
ThePrimeClock6132
🟠 redditUnofficial Jev plugin for coding agents: best practices, an API reference, and 150+ community projects. Evals included.
artificial
aaddrick12
🟧 hnEikos - OSS Jev-like modelemersonrsantos10
🟠 redditShould you let JEV make all of your decisions?
LocalLLaMA
tenkei_01026
🟧 hnShow HN: Matching Jev on BANKING77 at a thousandth of the costKasianFranks30
🟠 reddit10 Technical Questions About Jev
LocalLLaMA
Prashant-Lakhera02
🟠 redditI built a CLI for Jev-style typed decisions that can also run with local models
LocalLLaMA
muthuishere210100
🟠 redditLaya: replace LLM-as-a-judge with a 322M-parameter decision engine (26,639 stars in 9 days, hands-on test)
LocalLLaMA
AIFrontierReads04
🟠 redditQwen company already rushed out a Jev competitor. No open weights yet.
LocalLLaMA
pneuny8760
🟧 hnHow to build a Jev-style classifier with DiffusionGemma and vLLMteleforce11
🟧 hnShow HN: Jevpipe – a fast System 1 for AI agents, as a Unix pipebothlabs23
🟠 redditImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench
LocalLLaMA
Educational-Care786719861
🟧 hnShow HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMspragmatwice30
🟧 hnShow HN: Decide – Jev decisions in the shell, scripts, and agent skillsvsekhar20
🟠 redditGevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring
LocalLLaMA
inawhole08
🟠 redditI built a Unix pipe that lets Claude Code use Jev as a fast System 1
ClaudeAI
bothlabs08
🟧 hnWe swapped our LLMs for Jev. It's 39% cheapergczh74
🟠 redditTrained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)
LocalLLaMA
Usual_Maximum76733611
🟧 hnJeff – Jev-compatible 0.8B decision models, trained at home, ~30 msfirelex572217
🟧 hnClassify 6700 pages for $1 with a Jev-compatible VLM APIfzysingularity20
🟠 redditI routed Claude Code tasks with a cheap decision model. A cost saving is not established.
ClaudeAI
Monglong_korea11
🟧 hnStanford and Nvidia's open Jev-like modeloscarfr10
🟧 hnJeeves. Reasoning improves Jev-like decision modelsnicowaltz24294
🟧 hnWhat TypeSafe Got Right with the Jev Launchnkko20
🟧 hnShow HN: Jevstiller – Distill Jev into a local model, with a disagreement boundtgluck6515
🟧 hnJev-compatible document classification with open-weight VLMsfzysingularity10
🟠 redditDecisions API
OpenAI
Valuestudent34
🟠 redditTypeNotSafe
singularity
Technical-Will-286245881
🟧 hnOpenJev: An open-source, Jev-compatible System One decision enginerdudekul20
🟧 hnOpenAI Launches Decisions APIdvrp62
🟧 hnOpenAI Answers TypeSafe's Jev with a Decision API Built on Lunayawnxyz61
🟠 redditJev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included
LocalLLaMA
One_Temperature598305
🟠 redditI cut cost and latency on my search engine with Jev [D]
MachineLearning
AccomplishedEvent27310
🟧 hnOpenAI Launches Decisions API for Fast AI Choices
Retrieved article excerpt

Open article · Retrieved 2026-09-30T05:26:16.827020+00:00

[@OpenAIDevs](https://x.com/OpenAIDevs)

[OpenAI Developers](https://x.com/OpenAIDevs)

[OpenAI](https://x.com/OpenAI)

[@OpenAIDevs](https://x.com/OpenAIDevs)

[10h](https://x.com/OpenAIDevs/status/2105003318917697873)

Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna.
Define questions and possible answers to classify content, route requests, or choose an agent’s next action.
Available in limited preview.

![](https://pbs.twimg.com/media/HTZ6Dv2bAAAEavG?format=webp&name=medium)

00:00

[170](https://x.com/OpenAIDevs/status/2105003318917697873)

333

4.6K

451K
soltanov24
🟧 hnUsing Jev as a Search Rerankerdustincoates10
🟧 hnShow HN: Halv cut AI agent cost by 57.1% using Jevvillaspedro20
🟧 hnLaya: Multilingual, non-autoregressive System 1 decision enginehandfuloflight10
🟧 hnJev, Checkedwilsonwu810
🟧 hnShow HN: Snap – local first decision engine open weightemnlmn10
🟠 redditBenchmarking small confidence scoring decision (Jev, Laya) models [P]
MachineLearning
iam_gkrishna10
🟠 redditNIRNAY: 450M decision model beats Jev on Banking77, runs on CPU
LocalLLaMA
BrilliantSecret14307
🟠 redditnew system 1.5 model beats jev by a long shot on all 4 baselines and is #1 on image jev bench (yes its multi modal)
singularity
boneMechBoy694203540
🟧 hnClef: Open-source decision models, and new RL fine-tuning platformjasondavies626217
🟠 redditnew system 1.5 model beats jev by a long shot on all 4 baselines and is #1 on image jev bench (yes its multi modal)
artificial
boneMechBoy6942034
🟠 redditGuys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)
LocalLLaMA
Altruistic_Heat_9531012
🟧 hnDecision models like Jev don't beat LLM-as-a-judge or traditional classifierstomncooper13549
🟧 hnHow accurately calibrated is Jev?dblack1270563
🟠 redditBuilt an open-source tool that applies to internships for me
ClaudeAI
q3ndi23
🟠 redditJev: Not Frontier, But Still Worth Your Attention
LocalLLaMA
enn_nafnlaus05

Interpretation history

Decision trace