Jev is the first 'System One' model from TypeSafe AI, a San Francisco startup founded in 2024 and led by ex-OpenAI's Diogo Almeida, which emerged from stealth on 15 September 2026 alongside a $40M seed round led by DCVC at a reported ~$200M valuation. It is a proprietary model that replaces free-form text generation with typed probabilistic answers (choice/score/yes-no plus confidence values) that software can branch on directly; TypeSafe claims frontier-comparable quality at 70–500ms end-to-end and 40–200x faster / 40–400x cheaper than autoregressive LLMs, trained on synthetic data via a proprietary RLCD method, with no architecture, weights, or technical paper published — though the org's own GitHub fork of the LLaDA diffusion-LLM repo is consistent with outside observers' suggestion that it wraps an open-weight base, and its docs confirm text-only native input and shared weights with no per-customer fine-tuning. The supplied web material is entirely launch-side (vendor blog and docs, Business Wire release, Wikipedia stub, Forbes valuation piece) and carries none of the independent record the case has accumulated: third parties confirm bounded accuracy plus a real short-request cost/latency edge, but the load-bearing calibrated-probability claim is refuted by multiple independent ground-truth tests (jevals, a 400-roll fair-die audit, institutional benchmarks), the open-alternative field is saturated at 25+ models, and OpenAI answered the category within two weeks with a Luna-based Decisions API whose pricing, calibration approach, and image-input scope remain unverified.
| source | object | author | score | comments |
| 🟧 hn | Jev: New frontier model 40-400x cheaper and 20-200x fasterRetrieved article excerptOpen article · Retrieved 2026-09-15T21:23:17.811922+00:00 TypeSafe announces System One models and Jev
[Read More](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
[TypeSafe AI](https://typesafe.ai/)
[Manifesto](https://typesafe.ai/manifesto)
[Our Team](https://typesafe.ai/team)
Join Waitlist
TypeSafe announces System One models and Jev
[Read More](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
[TypeSafe AI](https://typesafe.ai/)
[Manifesto](https://typesafe.ai/manifesto)
[Our Team](https://typesafe.ai/team)
[∵ Back](https://typesafe.ai/)
Company News
Sep 14, 2026
# Introducing System One Models and Jev
*Diogo Almeida, founder, TypeSafe*
Models have been superhuman at chat for years, so where is all the automation?
This has been my driving question for the last four years. At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. That work ended up as the research behind ChatGPT. At the time, I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing.
After two years in stealth, countless technical challenges, and research breakthroughs… I am beyond excited to announce that today, TypeSafe AI is releasing our first **System One Model**: a new class of frontier models built to make fast, structured decisions that software can use directly.
We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).
Our first public model is **Jev**, available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient. While Jev gives up string generation, it’s optimized for structured outputs and *can’t* hallucinate.
Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.
Extraordinary claims require extraordinary evidence so see below for the receipts. 💅
## Frontiers, Old and New
| | **Existing LLMs** | **System One + Jev** |
| --- | --- | --- |
| Optimized with | Reinforcement Learning with Human Feedback (RLHF) / Reinforcement Learning with Verifiable Rewards (RLVR) | Reinforcement Learning for Calibrated Decisions (RLCD) |
| Optimizes for | Human preference: writeups and chat responses that human raters prefer. Verifiable rewards: outputs that can be programmatically verified. | Calibrated decisions: answers with epistemically honest probabilities on System One tasks. |
| Inputs | Unstructured data (e.g. text) with an emphasis on **sequential messages**. | Unstructured data (e.g. text) with an emphasis on **structured program state**. |
| Outputs | **Strings / generated text.** Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails. | **Type-safe structured values.** Possible outputs and structure are [defined in advance](https://docs.typesafe.ai/). The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores. |
| Sampling | **Sequential.** Generates one token at a time, each conditioned on the last. | **Parallel.** Generates all outputs in a single query. Incredibly efficient and hardware-aware. |
| Cost | Input tokens: from $0.20 to $10 / MTok. Output tokens: ~5x more expensive than input tokens. | Input tokens: $0.042 / MTok ($42 per billion tokens). Output tokens: FREE (too cheap to meter). |
| Speed | **End-to-end response time is** [**3 to 329 seconds**](https://llm-benchmarks.diegoromero.es/)for frontier models. Fast enough for interfacing with humans, but a big bottleneck when integrated in code. | **End-to-end response time is 70ms-500ms** for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries. |
| Confidence | Even if prompted for a confidence estimate, models tend to be overconfident and inconsistent. If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task. | Always communicates confidence and uncertainty with every output. Calibrated: higher confidence means higher accuracy. More consistent: returns similar answers for similar inputs. |
| Use cases | **Human-in-the-loop tasks (chatbots, copilots, coding agents).** General and powerful, but requires human oversight because their freedom also means they might go off the rails. **Verifiable problems (math proofs, kernel optimization).** When correctness can be checked cheaply and automatically, LLMs can generate, test, and iterate until they find something that works. **Demos.** The flexibility of strings allows it to be incredible for quickly making prototypes that only work sometimes. | **AI-Powered Workflows / smart if-statements.** Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems. **Map-reducing** **over big data.** Turn petabytes of data into features and insights. **Real-time applications.** 100ms speeds means you can use AI in your applications where UX is critical.**Verify everything.** Score, judge, verify, guardrail, and detect jailbreaks of LLM prompts, reasoning traces, and/or outputs. |
##
## Evidence / Technical Results
We love skeptics, and are skeptics ourselves.
There are some claims you can easily verify:
- **Speed per call:** We truly are that fast, though our published evals are generally run from our laptops on the West Coast (this is where our service is currently based).
- **Cost per call:** We make our pricing transparent. We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).
- **No type errors**: This would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible.
For our bolder claims, we want to provide as much nuance as we can.
### Side-by-side demonstration
Our side-by-side demo shows a key difference between our models and LLMs: Jev outputs all probabilities in parallel instead of autoregressively generating by token. Strings are extremely powerful and general, but costly. “Giving up” strings actually gives us a lot of superpowers!
[](https://framerusercontent.com/assets/g4FBJpKG573R4zYuoY3HLJIMCu4.mp4)
Nuance:
- For people with early access to TypeSafe, here is the [actual query](https://console.typesafe.ai/playground?share=shr_13a74b495fb786c4bd7964f11597301e7c9).
- The query is highly simplified and `questions` were chosen to have descriptive, human-readable keys so that the output on the screen is understandable.
- The `state` is also a short, dense, and detailed paragraph, to emphasize the difference in sampling methodology. The relatively shorter input paints our model in an advantageous light.
- For the keen eyed, for the recorded run, the only disagreement with GPT-5.6 Terra is on “Churn likelihood level”. The actual answer seems genuinely ambiguous to us.
- We used GPT-5.6 Terra with default reasoning for this example, because we’ve found it to be the most comparable at intelligence to Jev on average.
- Fun fact: a similar demo was what convinced us to go all-in in the direction of System One Models!
### Workflow evals
We made a new type of evaluation to measure how well AI works within code. We don’t optimize for a ground truth classification orand allow the harness and model to change (potentially allowing for overfitting via harness engineering). Instead, we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities.
Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).
Jev is off the charts – owning the Pareto frontier for almost 2 orders of magnitude. We also compare to models with a generated prompt doing all the logic in their chain-of-thought, but this tends to do significantly worse than using the workflow itself.
Note that the calls here are significantly more complex than the side-by-side demonstration above. That’s because they’re more representative of the types of production workloads needed for true business automation. Below is the simplest of the 4 workflows we’re publishing:
The most reliable real-world workflows tend to have many independent, decomposed questions, with fine-grained behavior that’s dependent on probabilities instead of discrete decisions. The end result is discrete branching, but how we get to a final answer involves a lot of domain-specific engineering that needs to be done highly consistently.
See [our workflow evals site](https://evals.typesafe.ai/) for all the details: examples, disagreements, full queries, and each workflow.
Nuance:
- This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.
- These content of these workflows were not deliberately chosen nor constructed to make our model look good, and are not in our training distribution. However, they were made by individuals on our model capabilities team, so some bias could exist.
- We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models.
- The LLMs use our [System One LLM](https://github.com/typesafe-ai/system-one-adapter-python) wrapper, which constrains LLMs to output structured decisions compatible with our API. We have found this to be the most accurate way to get decisions from LLMs, but this tends to be slower and more expensive than giving decisions without probabilities.
### Hallucination and Type-safety
Hallucination and type-safety are intrinsically related, and we think the latter is table stakes for automation. Having a hallucinated tool call is inconvenient in an agent, but is an absolute deal-breaker if it’s part of a system with latency guarantees or it’s buried several layers deep in a dependency chain. Existing models, *no matter how smart*, still hallucinate and have type errors.
Nuance:
- The numbers for LLMs are from OpenRouter i.e., there almost certainly is bias here: more complex queries might be routed to better models.
- Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.
### Fun Demos
Perhaps the most exciting part of our work is enabling new use cases. We have a lot more to show you, but here are a couple of the team’s favorites:
#### Doom
We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.
[](https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvjpY.mp4)
Nuance:
- The demo is on structured state as a data structure with text, not on images (yet…)
- A non-AI doom bot could play better, but we wanted a bot that was reactive to different representations of game state, and most importantly… following instructions was cool as heck!
#### Wikiracing
The objective of the game is to start on one Wikipedia page and reach a specific ot | albelfio | 1908 | 496 |
| 🟧 echo.blog ⭐ | Announces Jev in early access with parallel typed probabilistic outputs, $0.042 per million input tokens, and workflow evaluations reporting | Diogo Almeida, TypeSafe AI | — | — |
| 🟧 hn | GLiClass: Open-Source JEV | handfuloflight | 2 | 1 |
| 🟠 reddit | TypeSafe AI releases AI model called Jev. Rather than generating text, it makes decisions. Its hallucination rate is far lower and its outputs are very cheap compared to traditional LLMs. singularity | Profanion | 157 | 108 |
| 🟠 reddit | Near Here got early access to TypeSafe Jev, so we tested it for local event validation, tuning each model’s prompt individually. In our tests, Jev delivered up to 5.7× faster responses, 98% lower cost and 12 percentage points higher accuracy - see the results, methodology and limitations artificial | jon_reed | 2 | 0 |
| 🟠 reddit | LocalJev? LocalLLaMA | SomewhereAtWork | 77 | 38 |
| 🟠 reddit | New type of LLM released today " it’s optimized for structured outputs and can’t hallucinate" singularity | Oh_boy90 | 15 | 24 |
| 🟠 reddit | Qwen3.5 4B + grabbing logits is almost "Jev"? Or even just Qwen Reranker? LocalLLaMA | theoleecj_n | 49 | 14 |
| 🟧 hn | Reverse-engineered Jev-like model | rochansinha | 160 | 23 |
| 🟧 hn | vLLM: Jev-like mode for the DiffusionGemma model | mmastrac | 1 | 0 |
| 🟠 reddit | I built jev architecture and model one year back for sales LocalLLaMA | Nandakishor_ml | 3 | 1 |
| 🟠 reddit | I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper LocalLLaMA | Nandakishor_ml | 1544 | 133 |
| 🟧 hn | Jev Ultrafast: A browser agent with a dynamic, indexed action space | rahimnathwani | 60 | 4 |
| 🟠 reddit | I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper LocalLLaMA | Nandakishor_ml | 3351 | 318 |
| 🟧 hn | Open-sourced jev architecture last year with model,paper and dataset | nandakishor_ml | 39 | 9 |
| 🟠 reddit | If you have access to Typesafe, try this out singularity | purealgo | 4 | 3 |
| 🟧 hn | Open-jev: One-pass option scoring with Gemma 3 4B, similar to jev | suriyaG | 4 | 1 |
| 🟠 reddit | Jev from TypeSafe.ai is getting hyped quite a bit on X. Lots of fun use cases. Not a LLM but a super fast/cheap decision engine with Luna-level intelligence singularity | manubfr | 45 | 19 |
| 🟧 hn | AI Startup launches a faster and cheaper alternative to LLMs for AI automation | Sarvaturi | 1 | 0 |
| 🟧 hn | Jev for Home Assistant | AboveColin | 1 | 0 |
| 🟠 reddit | Jev is amazing! I'm letting it play Pokemon Red with a harness being built by Opus 5 in real-time — follow along! artificial | Boydbme | 0 | 0 |
| 🟧 hn | Jevmlx | trollied | 2 | 0 |
| 🟧 hn | Benchmark Jev vs. Gemini Flash and Claude Fable on Code Review | gemanor | 2 | 0 |
| 🟧 hn | Show HN: Sokit – a LangChain like harness for Jev (or other System 1 models) | phantomCupcake | 2 | 0 |
| 🟠 reddit | What do you think about Jev and RLCD in general? (Here is my personal take) artificial | Haghiri75 | 1 | 7 |
| 🟠 reddit | I thought I'd found a model 5000x cheaper than Claude for filling web forms. My benchmark was wrong. ClaudeAI | imaxalpha | 0 | 3 |
| 🟧 hn | Using Jev for Claude Code model routing | hassleblad23 | 2 | 0 |
| 🟠 reddit | jev reproductions tracker. keeping up with jev reproduction efforts LocalLLaMA | apolinariosteps | 22 | 3 |
| 🟠 reddit | Playing Doom using Jev by System One, the hottest new thing singularity | ascii_heart_ | 2 | 1 |
| 🟧 hn | Show HN: Jev routing coding tasks to Grok Build or Codex Astra | joshcsimmons | 1 | 0 |
| 🟧 hn | Is-odd-jev – check if a number is odd, with a calibrated probability | alxcrt | 1 | 0 |
| 🟧 hn | Talk to JEV | mkotlikov | 1 | 0 |
| 🟧 hn | Mini-Jev – typesafe's Jev implemented on top of an LLM locally | phyrex | 3 | 0 |
| 🟠 reddit | still doesn’t get what Jev is…..is it just a more generalised BERT? LocalLLaMA | AdRepulsive7837 | 145 | 47 |
| 🟧 hn | Show HN: Open-Source Alternative to TypeSafe.ai | ptitov | 2 | 1 |
| 🟧 hn | Open alternative to TypeSafe's Jev, running locally on your own GPU | ikerM | 1 | 0 |
| 🟠 reddit | Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo LocalLLaMA | Nandakishor_ml | 771 | 130 |
| 🟠 reddit | New 'decision' model Jev (developed by co-creator of ChatGPT) is playing Subway Surfers in real-time singularity | Cagnazzo82 | 547 | 105 |
| 🟧 hn | Probably – a programming language for LLM workflows, powered by Jev | porridgeraisin | 1 | 0 |
| 🟠 reddit | Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it LocalLLaMA | Every-Comment5473 | 78 | 20 |
| 🟧 hn | OpenJev | ilreb | 681 | 284 |
| 🟠 reddit | I use beads with Claude Code, but have found mass labelling of issues inefficient and costly. The new Jev model solves this problem, so I made a tool which uses it to label your beads. ClaudeAI | bobo-the-merciful | 1 | 4 |
| 🟧 hn | Ulka: Browser agent powered by Jev and fx.sh (experimental) | razaan | 1 | 0 |
| 🟠 reddit | I made a Claude Code skill that reaches for Jev before writing another regex, and writes down what happened ClaudeAI | bobo-the-merciful | 1 | 0 |
| 🟧 hn | Show HN: Jev vs. GPT-5.6 and Claude Haiku at Pong | matt_oriordan | 9 | 3 |
| 🟠 reddit | Jev vs Claude. i built the benchmark so you don't have to artificial | lutian | 0 | 1 |
| 🟠 reddit | Jev is amazing! I'm letting it play Pokemon Red with a harness being built by Opus 5 in real-time — follow along! singularity | Boydbme | 54 | 44 |
| 🟠 reddit | Jev playing 9 real-time classic games simultaneously with a single API call for $1.80/h singularity | manubfr | 127 | 46 |
| 🟧 hn | NanoJev | shock | 1 | 0 |
| 🟠 reddit | Diego Almeida, fondateur de Typesafe AI, présente JEV, artificial | Winter-Mix-5155 | 1 | 1 |
| 🟧 hn | Lmjtfy – Ask Jev a yes or no question | rbaudibert | 6 | 1 |
| 🟧 hn | TypeSafe / Jev latency-focused demos built by Devin | rahimnathwani | 2 | 0 |
| 🟧 hn | Show HN: Jeff – A read-only CLI for semantic code review using Jev | imalessandro | 3 | 0 |
| 🟧 hn | Jev as a Primitive Feature of Ruby | idz | 1 | 0 |
| 🟠 reddit | Routing between a 3B, a 4B, a 12B and a 26B MoE with Jev > regex LocalLLaMA | clduab11 | 15 | 7 |
| 🟧 hn | I used Jev to control a swarm of 15 simulated drones in real time | khordoo | 1 | 0 |
| 🟠 reddit | How is RLCD (jev) RL? [D] MachineLearning | Relative_Wallaby_823 | 20 | 19 |
| 🟧 hn | TypeSafe AI's Jev Is Not an LLM – and That May Be the Point | Bluestein | 2 | 0 |
| 🟠 reddit | Digit-logits-based classifier with llama.cpp LocalLLaMA | rhinodevil | 5 | 0 |
| 🟠 reddit | Used Claude to help Jev become a gamer ClaudeAI | oldmoldycake | 46 | 12 |
| 🟠 reddit | I benchmarked Jev aginst gpt-5.6-luna! OpenAI | LowNefariousness9966 | 125 | 42 |
| 🟧 hn | Laya the open source version of Jev | nandakishor_ml | 1326 | 305 |
| 🟠 reddit | Hundreds of examples of Jev use cases singularity | manubfr | 134 | 34 |
| 🟧 hn | kev: Jev-like model built on Qwen2.5-0.5B | tosh | 1 | 0 |
| 🟧 hn | A local Jev backed by DiffusionGemma | awei | 1 | 0 |
| 🟠 reddit | Von: Open-source 395M "System One" model LocalLLaMA | wFXx | 182 | 77 |
| 🟠 reddit | Jev Cuts AI Decision Costs 100x And Vercel, Cloudflare Rushed To Add It singularity | Prudent-Sorbet-5202 | 346 | 80 |
| 🟧 hn | Show HN: CUA-S1 – A System One Model for Computer Use | frabonacci | 83 | 8 |
| 🟧 hn | Show HN: VisionLaya: Jev with Vision capabilities | someguy101010 | 1 | 0 |
| 🟠 reddit | Jev vs Luna OpenAI | aniketmaurya | 0 | 3 |
| 🟧 hn | In 2024 I fine-tuned an LLM. Jev could have removed the side quests | jc4p | 2 | 1 |
| 🟠 reddit | TypeSafe AI is what the software industry has been waiting for!!! LocalLLaMA | BoyInDaBox89 | 0 | 10 |
| 🟧 hn | Show HN: S1Code, a decision-first Rust coding agent with Jev | merthdotxyz | 1 | 0 |
| 🟧 hn | Show HN: Jev to JSON-Schema | scosman | 1 | 0 |
| 🟧 hn | Testing Jev as a validation gate for drug-discovery agents | fred_tandemai | 3 | 0 |
| 🟧 hn | Show HN: Jev, Fly Me to the Moon | rahmanyoo | 2 | 0 |
| 🟧 hn | Show HN: Jev-align, a CLI to calibrate Jev to your judgement | sethkim | 2 | 0 |
| 🟠 reddit | I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom LocalLLaMA | shniydder | 217 | 98 |
| 🟠 reddit | Jev demos everywhere — I open-sourced the boring agent-loop recipes for Codex / Claude Code / OpenCode ClaudeAI | OscarwhDs | 1 | 1 |
| 🟠 reddit | AI Plays Streetfighter 2 In Real Time OpenAI | Smartaces | 39 | 38 |
| 🟧 hn | Jev vs. classical ML. Strong on sentiment: Mixed across tasks | theanonymousone | 2 | 0 |
| 🟧 hn | Jev is the fastest-adopted model in AI Gateway history | flashbrew | 2 | 1 |
| 🟠 reddit | a local Jev-style decision head onto Qwen 2.5 1.5B LocalLLaMA | CryOrganic8886 | 16 | 2 |
| 🟠 reddit | Jev benchmark uses Fable as truth but scores vs Opus ClaudeAI | Successful-Farm5339 | 0 | 1 |
| 🟧 hn | A MySQL plugin that filters rows by meaning (built on TypeSafe Jev) | maayanlevy-hn | 2 | 0 |
| 🟧 hn | DuckDB extension: typed Jev answers as real SQL types | jmrothweiler | 2 | 0 |
| 🟠 reddit | Is Typesafe based/derived from work done by the Laya author? LocalLLaMA | ECrispy | 31 | 30 |
| 🟠 reddit | What is JEV and what is it used for? LocalLLaMA | Hot_Example_4456 | 420 | 305 |
| 🟠 reddit | Dethrone - a PoC game testing Jev from typesafe.ai, the model that generates decisions instead of text. Evals at the Chronical link at bottom. artificial | aaddrick | 1 | 0 |
| 🟧 hn | Show HN: lgtm? – Jev-powered checks that make agents test | jennmueng | 1 | 0 |
| 🟧 hn | Auto approve pull requests with Jev | infiniteregrets | 2 | 0 |
| 🟧 hn | Show HN: A minimal Pareto-optimal OpenRouter model router for pi, based on Jev | 7777777phil | 2 | 0 |
| 🟠 reddit | Typesafe's JEV model work as an LLM [P] MachineLearning | dwarfLevi | 6 | 1 |
| 🟧 hn | Laya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second) | putna | 150 | 30 |
| 🟧 hn | Jev compiler – Turn the rules your agent keeps ignoring into gates you can test | doronp | 2 | 0 |
| 🟧 hn | Show HN: Real Jev decisions on a simulated robot fleet – $24.57 per million | chorylee | 2 | 0 |
| 🟠 reddit | laya.cpp: Optimized laya near-instant decision making LocalLLaMA | lkarlslund | 90 | 26 |
| 🟠 reddit | A Jev-style model fine-tuned on Qwen3.5 4B LocalLLaMA | nato_nob | 31 | 22 |
| 🟧 hn | Using Jev as a teacher to help an SLM write better stories | nutanc | 1 | 0 |
| 🟠 reddit | Jev as a classifier is ok but not the part thats important LocalLLaMA | 2BucChuck | 0 | 6 |
| 🟧 hn | I turned Jev into a (lousy) chatbot | kp1197 | 169 | 48 |
| 🟧 hn | Someone made jev play Atari games | thomask1995 | 3 | 0 |
| 🟧 hn | Show HN: Testing a non-generative decision model on 5,500 CLINC150 inputs | tgdhtdujeytd | 3 | 0 |
| 🟠 reddit | DIY Jev LocalLLaMA | Malfeitor1235 | 75 | 32 |
| 🟧 hn | Show HN: jevals – replacing LLM judges with typed Jev decisions | gbayomi | 28 | 1 |
| 🟧 hn | Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev | 0x1997 | 2 | 0 |
| 🟧 hn | Jev – System-1 Agent Architecture Radar (open-source) | noobplus | 2 | 0 |
| 🟠 reddit | convaiinnovations/laya (multilingual, non-autoregressive System 1 decision model) LocalLLaMA | Balance- | 28 | 1 |
| 🟧 hn | Jev, Prolog, Pi, and the dream of probabilistic logic programming | schmuhblaster | 16 | 0 |
| 🟠 reddit | I built JevGraph, an open-source pipeline to turn documents into evidence-backed knowledge graphs LocalLLaMA | richie9830 | 0 | 7 |
| 🟧 hn | Find-jevable-code – audit a repo for Jev-replaceable decisions | ss_y2n | 2 | 0 |
| 🟧 hn | Kev: Tiny Jev-like family of decision models built on top of Qwen3.5 | tosh | 422 | 191 |
| 🟧 hn | Models Watching Models | tosh | 1 | 0 |
| 🟧 hn | Show HN: Jeeva – A modular trading engine for mid-frequency trading using Jev | mugiwaraa_eth | 1 | 0 |
| 🟧 hn | Show HN: Grade text from the CLI with custom rulesets and Jev | lukstei | 3 | 0 |
| 🟧 hn | Jev and AI SDK Template by Vercel Labs | flashbrew | 1 | 0 |
| 🟠 reddit | Where Jev can take work off Claude and where it cannot, from TypeSafe's own docs ClaudeAI | prakersh | 1 | 1 |
| 🟠 reddit | A $40M model is being sold on calibrated confidence, and no calibration data has been published artificial | prakersh | 9 | 11 |
| 🟧 hn | Show HN: Jev-CLI – CLI wrapper for JEV typesafe AI model | joshLong145 | 2 | 0 |
| 🟧 hn | I made an LLM using 521 Jev models | skillseeddev | 1 | 1 |
| 🟧 hn | We Tested Jev on 100 Agent Tool Calls | arseny_info | 11 | 0 |
| 🟧 hn | Show HN: OpenDecision – a 400M zero-shot model makes local decisions, plays DoomRetrieved article excerptOpen article · Retrieved 2026-09-21T15:25:13.829826+00:00 # OpenDecision
OpenDecision answers typed questions about application state and documents. It runs a local natural language inference model and returns structured values.
## Doom demo
OpenDecision chooses actions for a bot in ViZDoom's Deadly Corridor. This is the Skill 5 recording.
[
Your browser does not support embedded video. [Open the recording](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill5.mp4).
](https://cdn.jsdelivr.net/gh/deepanwadhwa/OpenDecision@main/demos/doom/opendecision-doom-skill5.mp4)
[Watch Skill 1](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill1.mp4) | [Watch Skill 3](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/opendecision-doom-skill3.mp4) | [Run the demo](https://github.com/deepanwadhwa/OpenDecision/blob/main/demos/doom/README.md)
## Install
With `pip`:
```
pip install OpenDecision
```
With `uv`:
```
uv add OpenDecision
```
[Continue to the get started guide](https://deepanwadhwa.github.io/OpenDecision/quickstart/)
## What it provides
| Type | Use | Result |
| --- | --- | --- |
| `Choice` | Select one option. | Option name and probabilities |
| `Noul` | Test one statement. | A score from 0 to 1 |
| `Score` | Use an ordered scale. | Weighted score and probabilities |
| `Relation` | Compare evidence with a statement and its opposite. | `supports`, `contradicts`, `unknown`, or `conflicted` |
| Document decisions | Ask questions about text or JSON. | Answers and source passages |
[See code examples for each primitive](https://deepanwadhwa.github.io/OpenDecision/primitives/)
## Start here
- [Get started](https://deepanwadhwa.github.io/OpenDecision/quickstart/): install the package, run a Python example, and start the API.
- [Primitives](https://deepanwadhwa.github.io/OpenDecision/primitives/): use `Choice`, `Noul`, `Score`, and `Relation`.
- [Document decisions](https://deepanwadhwa.github.io/OpenDecision/document-decisions/): ask questions about long text or JSON and select a yes/no mode.
- [Evidence and rules](https://deepanwadhwa.github.io/OpenDecision/evidence-and-rules/): rank evidence and combine facts with rules.
- [Examples](https://deepanwadhwa.github.io/OpenDecision/examples/): run the Doom demo and review the insurance and GDPR examples.
## Interfaces
| Interface | Use |
| --- | --- |
| Python | Call OpenDecision in the same process as the application. |
| `POST /v1/systemone` | Send state and typed questions to the API. |
| `POST /v1/documents/decide` | Send a document and typed questions to the API. |
| TypeSafe-compatible endpoint | Use a compatible TypeSafe SDK client with a local server. |
## Basic example
```
from opendecision import OpenDecisionEngine
engine = OpenDecisionEngine()
result = engine.choice(
state="The customer was charged twice.",
instructions="Which team should handle this request?",
criteria={
"billing": "Payments, invoices, refunds, and duplicate charges",
"technical": "Software bugs",
"sales": "Pricing and purchases",
},
)
print(result["choice"])
# billing
``` | dwa3592 | 6 | 0 |
| 🟧 hn | Show HN: Run Jev-style models locally on Mac with 0.74 GB RAMRetrieved article excerptOpen article · Retrieved 2026-09-21T15:25:28.263414+00:00 # Laya MPS
**Run Jev-style typed decisions locally on your Mac with low RAM usage and fast responses.**
Laya delivers typed decisions with ~32 ms median latency using ~2.1 GiB RAM on
M5 Pro, with a slower ~0.74 GiB mode for lower memory use.
[Laya MPS Pong demo with live response latency](https://github.com/afshinm/laya-mps/blob/main/.github/assets/pong-demo.gif)
MPS stands for [Metal Performance Shaders](https://developer.apple.com/metal/pytorch/),
which PyTorch uses to run Laya on your Mac's GPU.
[Laya typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions)
chooses options, scores inputs, and estimates whether statements are true. The
English model specializes in customer service, invoices, security incidents,
and agent traces. It is not a general-purpose language model.
[Model comparison and benchmarks](https://github.com/afshinm/laya-mps/blob/main/benchmarks/README.md).
## Requirements
- Apple Silicon Mac (M1 or newer), macOS 14+.
- Git and [uv](https://docs.astral.sh/uv/getting-started/installation/). uv installs Python 3.12 if needed.
- Allow 4 GB of free disk space for the runtime, download cache, and ~843 MB model.
## Get started
Clone the repository and start the server:
```
git clone https://github.com/afshinm/laya-mps.git
cd laya-mps
./scripts/serve.sh
```
The first run installs dependencies and downloads the model. Open
**[the demo](http://127.0.0.1:8000/demo/)** and click **Play** or **Run Benchmark**.
The benchmark runs for 60 seconds and plots response latency.
The server runs on `127.0.0.1:8000`. Later runs use local files and inference
works offline. Stop it with **Ctrl+C**. To update a clone, stop the server,
run `git pull`, then run the script again.
## Memory settings
All settings run the same complete model in FP32. Measured on an **M5 Pro,
24 GiB RAM**, macOS 26.4:
| Setting | Peak process RAM | Median decision latency |
| --- | --- | --- |
| `minimal` | **0.74 GiB** | 171 ms |
| `reduced` (default) | 2.11 GiB | **32 ms** |
| `full` | 2.30 GiB | 32 ms |
RAM is the peak across the benchmark suite, excluding macOS, the browser, and
other apps. Latency uses fixed Pong inputs and excludes HTTP. All settings
matched exactly on **260/260 decisions**, including probabilities.
[Full measurements](https://github.com/afshinm/laya-mps/blob/main/benchmarks/README.md#memory-and-latency).
The default keeps transformer layers in RAM and reads embeddings from disk.
`minimal` also streams layers from disk; `full` keeps all weights in RAM.
To change the setting, stop the server and restart:
```
./scripts/serve.sh --memory minimal
```
How memory savings work
Resident weights load directly into their final FP32 allocation, avoiding a full
host-model copy. The checkpoint stores 16-bit values; expanding them to FP32
preserves their values. Computation stays in FP32 in every mode.
Disk embedding lookups fetch only the needed token rows, combine adjacent reads,
and restore the original token order. A 1 MiB row cache keeps the 196.75 MiB FP32
embedding table off the GPU in `minimal` and `reduced`.
`minimal` executes all 28 encoder layers and two decision-head layers through
shared buffers: 24.03 MiB on the host and 48.05 MiB on the GPU. Each forward
requests about 702.66 MiB of layer bytes. GPU work finishes before buffer reuse.
Multiple questions can require multiple forwards, which adds disk-read latency.
`F_NOCACHE` is requested on macOS, but OS caches may still serve reads; logical
read volume is not physical SSD traffic.
The runtime defaults `PYTORCH_MPS_LOW_WATERMARK_RATIO` to `0.00001` before GPU
allocation to encourage smaller Metal heaps and earlier reclamation. Explicit
environment values take precedence. This is a soft watermark, not a RAM limit;
the hard high watermark stays at its runtime default. When using Python directly,
create the engine before other MPS allocations. Effective settings appear in
`diagnostics=true` responses.
Implementation references: [Laya source](https://github.com/NandhaKishorM/laya/tree/6a5819129eb220570792e417e49723d697efd76f),
[MPS allocator](https://github.com/pytorch/pytorch/blob/08187d9e0fba026dc8217405802ab5381dc88d90/aten/src/ATen/mps/MPSAllocator.h),
[PyTorch settings](https://docs.pytorch.org/docs/2.14/mps_environment_variables.html),
[Safetensors format](https://github.com/huggingface/safetensors#format).
## Use the API
With the server running, choose a team for a support ticket:
```
curl --fail-with-body -sS http://127.0.0.1:8000/v1/decisions \
-H 'Content-Type: application/json' \
--data-raw '{
"state": "The export page returns HTTP 500. Our team cannot finish its work.",
"questions": {
"owner": {
"type": "choice",
"instructions": "Which team should investigate this report?",
"criteria": {
"billing": "Payments and invoices",
"engineering": "Software failures",
"unknown": "Insufficient information"
}
}
}
}'
```
Example response:
```
{
"model": "convaiinnovations/laya-typed-decisions",
"answers": {
"owner": {
"type": "choice",
"choice": "engineering",
"probabilities": {
"billing": 0.06904838234186172,
"engineering": 0.6391704082489014,
"unknown": 0.29178112745285034
},
"confidence": 0.24445859127951464
}
}
}
```
Check whether work is blocked and score the impact in one request:
```
curl --fail-with-body -sS http://127.0.0.1:8000/v1/decisions \
-H 'Content-Type: application/json' \
--data-raw '{
"state": "The export page returns HTTP 500. Our team cannot finish its work.",
"questions": {
"blocked": {
"type": "noul",
"instructions": "Does the report say that work cannot proceed?"
},
"severity": {
"type": "score",
"instructions": "Rate the operational impact described in the report.",
"levels": [
"Cosmetic issue with no work affected",
"Some inconvenience but work can proceed",
"Work is blocked for the whole team"
]
}
}
}'
```
Example response:
```
{
"model": "convaiinnovations/laya-typed-decisions",
"answers": {
"blocked": {
"type": "noul",
"noul": 0.7066293954849243
},
"severity": {
"type": "score",
"score": 1.786898910999298,
"probabilities": [
0.019407516345381737,
0.17428618669509888,
0.8063063621520996
],
"legend": {
"0": "Cosmetic issue with no work affected",
"1": "Some inconvenience but work can proceed",
"2": "Work is blocked for the whole team"
},
"confidence": 0.4951949684505006
}
}
}
```
Responses contain `model` and typed `answers`. `choice` selects a label, `noul`
is the probability of true (0–1), and `score` is a probability-weighted level
number (0–2 in this example). Confidence describes how concentrated the
probabilities are; it does not establish correctness. Evaluate the model on
your own inputs before relying on it.
Timing and diagnostics
Add `?metrics=true` to include engine evaluation time and current server-process
RAM. The demo measures full HTTP response time separately. Memory is physical
footprint on macOS, or RSS on CPU platforms without that measurement; it is
not peak RAM or whole-machine RAM.
```
curl --fail-with-body -sS 'http://127.0.0.1:8000/v1/decisions?metrics=true' \
-H 'Content-Type: application/json' \
--data-raw '{
"state": "The export page crashes. Our team cannot finish its work.",
"questions": {
"blocked": {
"type": "noul",
"instructions": "Does the report say that work cannot proceed?"
}
}
}'
```
Example response (timing and memory vary):
```
{
"model": "convaiinnovations/laya-typed-decisions",
"answers": {
"blocked": {
"type": "noul",
"noul": 0.6566364169120789
}
},
"metrics": {
"request_ms": 20.240125013515353,
"memory_bytes": 2068972792
}
}
```
Add `?diagnostics=true` for token counts, temperatures, logits, native
probabilities, and auxiliary action probability, plus runtime versions,
timing, memory snapshots, and storage I/O. Both query options can be combined.
Native Noul probabilities are ordered `[false, true]`; the public `noul` value
is the probability of true. The auxiliary action probability is a separate head.
API reference and limits
Base URL: `http://127.0.0.1:8000`. No API key is needed. The server accepts local
connections; cross-origin browser access is not enabled.
| Endpoint | Returns |
| --- | --- |
| `GET /health` | `{"status":"ready","busy":false}` |
| `GET /v1/config` | Model, revision, memory setting, device, precision, and limits |
| `POST /v1/decisions` | Model identifier and typed answers |
| `GET /docs` | Interactive API reference; UI assets load from a CDN |
| `GET /openapi.json` | API schema, available offline |
| `GET /demo/` | Pong and the latency benchmark; no CDN assets |
Send `Content-Type: application/json`. `state` accepts finite JSON and
`questions` contains 1–16 named questions. Optional top-level `instructions`
apply to every question. Questions run independently.
| Type | Input | Answer |
| --- | --- | --- |
| `choice` | `criteria`: 2–26 labels with descriptions | Label, probabilities, confidence |
| `score` | `levels`: 2–10 descriptions, low to high | Weighted zero-based score, probabilities, legend, confidence |
| `noul` | Optional `criteria` with `false` and/or `true` descriptions | Probability of true |
Instructions and descriptions are strings. Score probabilities follow level
order. The default model allows 1,024 formatted tokens per question; the optional
`english` checkpoint allows 512. Formatting includes the state, instructions,
and options. Inputs that would be shortened are rejected.
| Status | Meaning |
| --- | --- |
| `400` | Invalid host, encoding, or unparseable body |
| `413` | Body exceeds the default 1 MiB limit |
| `422` | Invalid fields, options, JSON, or context overflow |
| `503` | Inference is busy or model execution failed |
One inference runs at a time. A busy response includes `Retry-After: 1`; health
and configuration remain available. Runtime settings are chosen at startup.
The `X-Laya-MPS-Config` header identifies the active configuration on config and
decision responses, allowing the demo to detect changes during a run.
The answer types follow [Jev's typed decisions](https://docs.typesafe.ai/introduction/quickstart).
This API has its own request limits, Score `levels` field, and probability-list
format; it is not a drop-in Jev endpoint.
## CLI and setup
Save the JSON payload from a curl example as `request.json` to evaluate it
without starting a server:
```
uv run --locked laya-mps decide request.json --download
uv run --locked laya-mps decide request.json --metrics
```
Pass server options to `scripts/serve.sh`. Use `--port 8001` if port 8000 is busy,
`--checkpoint english` to evaluate the original English model, or
`--model-dir PATH` to use another model directory. `--device cpu` is an explicit fallback;
the published performance numbers use the Mac GPU. The default
`--question-batch-size 1` limits peak RAM. List all options with
`uv run --locked laya-mps serve --help`.
Setup checks and offline startup
```
uv run --locked laya-mps doctor
uv run --locked laya-mps setup
UV_OFFLINE=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 ./scripts/serve.sh
```
`doctor` checks GPU availability and model readiness. `setup` downloads the
pinned model revision. Models live in `.models/` unless `--model-dir` is set;
the helper script always runs from the repository root. `serve --download`
repairs incomplete installations and reuses complete ones. The setup marker is
written only after all required files are present. There is no automatic model
or device fallback.
Use a normal terminal if a restricted environment cannot access Metal. Setup
requires internet access and Git; t | afshinmeh | 3 | 0 |
| 🟧 hn | JevEmon: Typed Decisions over GBA RAM to Walk Pokémon FireRedRetrieved article excerptOpen article · Retrieved 2026-09-21T15:25:29.148006+00:00 # JevEmon
Jev walks a real Pokémon FireRed ROM.
**Bring your own ROM, we don't ship it.**
Not a screenshot agent. Not a bot that mashes A. Each leg of the walk, the code reads the overworld out of RAM, works out every place you could actually go from here — a door, a path to the next route, a Pokémon Center if your party needs one — and hands that list to [Jev](https://docs.typesafe.ai/models.md). Jev picks a destination; the code paths there and presses the buttons. If a wild Pokémon interrupts the walk, Jev fights it out with the same kind of typed decision, then the journey re-plans from wherever the encounter left you.
More on how this is wired: [ARCHITECTURE.md](https://github.com/daniel4x/JevEmon/blob/main/ARCHITECTURE.md)
[Jev walking from the Player's House through Pallet Town to Viridian City](https://github.com/daniel4x/JevEmon/blob/main/docs/journey.gif)
## Status
Verified milestones so far: Jev can leave the Player's House, cross Pallet Town, deliver itself through Route 1 (fighting and winning any wild encounters along the way), and reach Viridian City. Further legs of the journey (Oak's Parcel, the first Gym) aren't built yet — the walk currently ends the run once it reaches Viridian City.
## Try it
Grab a [TypeSafe API key](https://typesafe.ai) and put it in `.env`.
```
cp .env.example .env
```
Drop `Pokemon - Fire Red Version (U) (V1.1).gba` next to this README. Requires a Mac with Apple Silicon.
```
brew install mgba ffmpeg uv
uv run python scripts/setup.py
uv run python -m jevemon --prepare
uv run python -m jevemon
```
That builds the checkpoint, starts the local server, and opens <http://127.0.0.1:8765> for you. Hit Go live and sit with it. Edit the lineup in the page. Speed it up. Scrub the VOD after. Everything runs through `uv` — no shell wrapper scripts, no separate install step.
One walk, no browser:
```
uv run python -m jevemon --run --speed 4
```
Other useful commands:
```
uv sync # just the Python environment
uv run python -m jevemon --check-journey # offline, no API calls
uv run python -m jevemon --port 9000 # serve on a different port
```
## What Jev sees
Not the screen.
A JSON snapshot of the overworld: where you are, your party's HP, and every destination the pathfinder already confirmed is reachable from here — each with how many times you've already visited it. During a wild encounter, the snapshot switches to the same battle state as any other fight: HP, types, moves, stats, abilities, PP, and only the legal moves and switches for that turn.
Jev answers with one choice and a probability for every option offered. If the answer is illegal, the run stops. There is no backup brain.
On the ROM, a switch is the Pokémon's personality value, not "slot 3." The party menu moves around. Using the old slot would send out the wrong one. That note lives in [AGENTS.md](https://github.com/daniel4x/JevEmon/blob/main/AGENTS.md).
Edit the starting lineup in `config.json` or the UI.
## This is a demo
The interesting part isn't the exact team or the exact route. It's that a typed choice over honest game state — no screenshots, no free-text prompting — is already enough to walk an overworld and handle whatever interrupts it.
Fork it. Push the journey further. Teach it a destination beyond Viridian City, or a whole different region.
Third-party dependencies and their licenses are in [NOTICE.md](https://github.com/daniel4x/JevEmon/blob/main/NOTICE.md).
## License
Copyright © 2026 Daniel Alfasi. Licensed under the GNU General Public License v3.0 — see [LICENSE](https://github.com/daniel4x/JevEmon/blob/main/LICENSE). | alfasiii | 1 | 0 |
| 🟠 reddit | A weekend with Jev made my coding agents up to 31% faster ClaudeAI | bartlomein | 0 | 10 |
| 🟧 hn | Using Jev as an LLM Linter for OMP (Pi) Agent | goulinkh | 1 | 0 |
| 🟧 hn | Using TypeSafe AI in Bug Bounty | lampysecurity | 1 | 0 |
| 🟠 reddit | On a small pilot, Gemma 4 E2B's verbalized confidence put 99.6% of its judgments on 0.00, 1.00 or 0.50, and a separate judge (Jev & Laya-type arch) ranked the same passages better, AUROC 0.926 vs 0.804, including the ones Gemma called certain LocalLLaMA | clduab11 | 0 | 13 |
| 🟠 reddit | Jev vs Laya head to head benchmark [D] MachineLearning | bobo-the-merciful | 1 | 0 |
| 🟠 reddit | LLMs can already mostly do what Jev does if you limit them to 1 token of output singularity | arkuto | 0 | 14 |
| 🟠 reddit | I went through 600+ Jev builds. the interesting part isn't the flashy demos artificial | Sarthak999gupta | 0 | 4 |
| 🟠 reddit | Exploring small LLMs as classifiers to rival Jev (sharing what I've found) LocalLLaMA | OneFanFare | 8 | 7 |
| 🟠 reddit | Jev's calibration was measured. The LLMs won [D] MachineLearning | frappuccinoCoin | 0 | 1 |
| 🟠 reddit | Using Jev to automatically moderate social media posts according to your site's rules. OpenAI | RealDannyhvv | 0 | 10 |
| 🟧 hn | Jev: System One Models for Prod, Not God – With Diogo Almeida, CEO, TypeSafe AI | swyx | 4 | 1 |
| 🟠 reddit | DeepSeek 4.1 Flash (as a System One model) vs Jev LocalLLaMA | frappuccinoCoin | 0 | 8 |
| 🟧 hn | Show HN: Jevopt: Making intelligent compiler optimisation decisions with Jev | ramneet_singh | 1 | 0 |
| 🟠 reddit | Kev: tiny Jev-like decision models (0.8B/4B/9B) on Qwen3.5 you can train and run locally - the 9B fits a 32GB Mac LocalLLaMA | khiladi1729 | 22 | 10 |
| 🟠 reddit | jevals: locally runnable evals for agents using Jev-style decisions LocalLLaMA | byebaybay | 1 | 0 |
| 🟠 reddit | Jev is very good. Your own data is better. LocalLLaMA | TrifleHopeful5418 | 0 | 14 |
| 🟧 hn | Laya vs Jev head-to-head on identical inputs | beckford | 1 | 0 |
| 🟠 reddit | Jev to autoselect Claude model - co-creator of chatgpt's project ClaudeAI | fsharpman | 9 | 12 |
| 🟧 hn | Jev introduces a new shape of LLM | benwerd | 42 | 13 |
| 🟧 hn | jev-router: route to the cheapest model in claude code for your task | saikatsg | 2 | 0 |
| 🟧 hn | TypeSafe AI Jev vs. GPT-6 Astra | flashbrew | 4 | 0 |
| 🟧 hn | Autoresearch on Jev: improving a Jev harness to beat Deekseek | scosman | 2 | 0 |
| 🟧 hn | Where Jev worked for us, and where it didn't | aozisik | 2 | 0 |
| 🟧 hn | Show HN: Blink – A high-performance Jev like decision model for C and WASM | marcobambini | 1 | 0 |
| 🟧 hn | Show HN: Gemma 3 4B as a typed decision function in Rust (47 ms/decision) | zozo123-IB-IL2 | 1 | 0 |
| 🟠 reddit | Qevi-2B: A Jev-style finetuned model for image classification LocalLLaMA | Taronyuuu | 0 | 0 |
| 🟧 hn | OpenAI is about to eat Jev's lunch – Arcturus Labs | JohnBerryman | 299 | 212 |
| 🟠 reddit | I reverse-engineered Jev and rebuilt it as open weights. It beats mine on all 7 benchmarks; mine beats it on my own task with 395 labels. LocalLLaMA | s1lv3rj1nx | 0 | 15 |
| 🟧 hn | I deleted slow, expensive LLM turns in coding agents with Jev | respectattentio | 1 | 0 |
| 🟧 hn | I built a linter/skill that ensures requests to jev are structured correctly | suraj_phanindra | 1 | 0 |
| 🟧 hn | We put Jev in production against a cross-encoder. Here are the numbers | dennispi | 2 | 0 |
| 🟧 hn | Go-System-One: Jev-Like Results in Pure Go and SIMD/PTX | rcarmo | 1 | 1 |
| 🟧 hn | Show HN: Ego-jev – 0.4s typed decisions for browser agents | ZephyrDeng | 1 | 0 |
| 🟧 hn | Calibrating Jev as a Code Reviewer | Lectem | 1 | 0 |
| 🟧 hn | JevBench, a reproducible benchmark for typed decision models | florianstandhar | 116 | 32 |
| 🟠 reddit | JevBench LocalLLaMA | openSourcerer9000 | 0 | 3 |
| 🟠 reddit | Does Jev reveal hidden sexist and racist tendencies in AI? artificial | No-Wishbone2391 | 0 | 20 |
| 🟧 hn | Faster and local Jev like model for Mac | mkagenius | 1 | 0 |
| 🟧 hn | TinyJev -Tiny Jev-style decision model that runs offline | aglaweankit | 2 | 0 |
| 🟠 reddit | Is jev, text embeddings but without chunks? (results included) LocalLLaMA | No_Afternoon_4260 | 3 | 8 |
| 🟧 hn | I built the Jev architecture one year ago and open-sourced it | htk | 19 | 1 |
| 🟠 reddit | I tested Semif (openjev) vs Von vs Jev LocalLLaMA | KingPinX | 1 | 9 |
| 🟠 reddit | Using Jev to evaluate LLMs, RAG, and agents - not just give them a score LocalLLaMA | Charming_Group_2950 | 0 | 0 |
| 🟠 reddit | stuntd: a local Jev-compatible server on Laya that learns from your own traffic (no API key needed) LocalLLaMA | Inevitable-Log5414 | 24 | 10 |
| 🟧 hn | Jevify skill – Gets your existing agents running on Jev | soupz011 | 2 | 0 |
| 🟧 hn | Jev in 25 Lines of Python | bashbjorn | 576 | 186 |
| 🟠 reddit | Jev in 25 lines of Python LocalLLaMA | johnnyApplePRNG | 269 | 101 |
| 🟠 reddit | A proxy that watches the yes/no and pick-one decisions your app asks an LLM for, then trains a local model to make them for free artificial | Inevitable-Log5414 | 1 | 4 |
| 🟠 reddit | Kev - an open-source System One decision engine that can compete with Jev, using a local 9B model ClaudeAI | Miserable_Extent8845 | 1 | 3 |
| 🟠 reddit | I turned Qwen3.8-27B Q2_64 + llama.cpp into a fully TypeSafe AI-compatible Jev-like system. OpenAI API still intact! World’s first Vision-enabled Jev-like model! <10 GB VRAM, 170 ms on an RTX 3090 and ~140 tok/s in chat. 76% vs. 88% Jev-1.13 Acc. on a diverse 22,000-request typed-decision benchmark LocalLLaMA | kyr0x0 | 0 | 29 |
| 🟠 reddit | Mods: can we do something about half the forum getting filled with these advertising posts for Jev? LocalLLaMA | Acrobatic_Stress1388 | 1182 | 240 |
| 🟧 hn | Beating Jev's accuracy, speed, and cost with open models | rob313 | 2 | 0 |
| 🟧 hn | Laya MPs Source | piqufoh | 1 | 0 |
| 🟧 hn | Jev does not play dice | someguy101010 | 1 | 0 |
| 🟧 hn | Jev Can't Be Calibrated | alexmolas | 9 | 14 |
| 🟧 hn | I Couldn't Build Jev at OpenAI – Diogo Almeida, TypeSafe Co-Founder and CEO [video] | ABS | 4 | 0 |
| 🟧 hn | Show HN: Open Code for Jev | phegler | 1 | 0 |
| 🟠 reddit | I made a Go adapter that gives local llama.cpp models a Jev-like decision API LocalLLaMA | wenyani | 0 | 19 |
| 🟠 reddit | This time I tested reflex vs SemIf vs Laya vs Von vs jev LocalLLaMA | KingPinX | 2 | 3 |
| 🟠 reddit | Jev isn't new tech. Its marketing targets people who think AI started with LLMs. LocalLLaMA | tiensss | 773 | 288 |
| 🟠 reddit | jevals: locally runnable evals for agents using Jev-style decisions LocalLLaMA | byebaybay | 0 | 4 |
| 🟧 hn | Jev vs. LLMs on 770 "Am I the Asshole?" posts | dchristopoulos | 16 | 2 |
| 🟧 hn | Jev is 13.6x faster, 2.7x cheaper than GPT Luna 6 | sjmaplesec | 2 | 2 |
| 🟠 reddit | JEV broke down 724 live ads from 37 brands in 40 seconds for $0.09 of tokens artificial | sibraan_ | 0 | 10 |
| 🟧 hn | Jev deserves hype but not the type its getting | yididev | 1 | 1 |
| 🟧 hn | I tried using Jev for judgment based guardrails | deepanshsaxena | 2 | 0 |
| 🟠 reddit | I made deepmoney a while back. My new side project: using the stock market as the label for a Jev-style news screener on Qwen3.8-27B LocalLLaMA | Fun_Water2230 | 0 | 2 |
| 🟠 reddit | JEV almost dead: CLM vs JEV LocalLLaMA | R_Duncan | 422 | 180 |
| 🟧 hn | Contrastive Language Model (CLM): An Ultra-Fast System One Model | piyushsthr | 2 | 0 |
| 🟠 reddit | Jev using my computer to make a linkedin post singularity | bGivenb | 15 | 23 |
| 🟧 hn | What Is RLCD? The Secret Behind Jev | tnspacetime | 63 | 8 |
| 🟧 hn | Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll | kantahayashi | 4 | 9 |
| 🟠 reddit | I built an open-weight alternative to Jev / TypeSafe - introducing OpenJudgement-4B (early preview) LocalLLaMA | bakatristan | 14 | 13 |
| 🟧 hn | Jev Against a Cross-Encoder | felineflock | 3 | 0 |
| 🟧 hn | Local JEV, fine tuned in < 30 mins on Mac Air | shmc | 1 | 0 |
| 🟧 hn | Show HN: Cbjev – typed decisions about text from one encoder pass | tomek7667 | 1 | 1 |
| 🟧 hn | JEV-Star: Low-Cost StarCraft II Control with Language-Model Planning | tndl | 1 | 0 |
| 🟠 reddit | Opus 5.5 co-designed and ran a benchmark and evaluation of jev-1.13 ClaudeAI | giulioc84 | 1 | 2 |
| 🟧 hn | Show HN: Jev-pilot – Jev picks Claude Code's effort, model and skill per prompt | akramovic | 2 | 0 |
| 🟧 hn | Valen: A multimodal decision model inspired by Jev | taylorfinley | 1 | 0 |
| 🟧 hn | Show HN: Jev-browse – browser sub-tasks for coding agents at ~1/3 the cost | danielnc | 1 | 0 |
| 🟧 hn | Jev and System One Models: Calibration Beats Accuracy | pansuriyakartik | 13 | 11 |
| 🟧 hn | Jev Based Code Review | namanbhulawat | 46 | 55 |
| 🟧 hn | Perch: Semantic Code Linting with Jev | handfuloflight | 1 | 0 |
| 🟧 hn | Show HN: Jevper – an LLM API client constrained to the Jev wire format | czl_my | 1 | 0 |
| 🟧 hn | Show HN: Knowledge Signal – A JEV-powered rubric assessment tool for study notes | paperplaneflyr | 2 | 0 |
| 🟠 reddit | Jev didn't beat Claude, it just made me realize how much I overuse Claude ClaudeAI | Intrepid_Truth8898 | 1 | 4 |
| 🟧 hn | Show HN: JevPertus – Jev-style option scoring on Apertus | jaluus | 2 | 0 |
| 🟧 hn | I made a linter using Jev for things a regular linter can't catch | czxtm | 2 | 0 |
| 🟧 hn | A Jevlike using an escalation model to route between a classifier and an LLM | yogthos | 2 | 0 |
| 🟧 hn | ReAnchor, a Jev Optimizer in DSPy | dbreunig | 1 | 0 |
| 🟧 hn | Benchmarking Jev, Laya, and five open models under three stresses | lovegreenlife | 3 | 0 |
| 🟧 hn | Show HN: Jev Plays Pokémon Red | pancomplex | 272 | 121 |
| 🟠 reddit | Jev vs. Kev: open-source Jev alternative tested side by side LocalLLaMA | facethef | 118 | 42 |
| 🟧 hn | Jev vs. Kev: open-source Jev alternative tested side by side | felix089 | 12 | 2 |
| 🟧 hn | Ollaya – Ollama for open-source, Jev-style decision models | Ardakilic | 578 | 139 |
| 🟧 hn | Contrastive Language Models: A Fast, Generalizable System One Model | yarapavan | 1 | 0 |
| 🟠 reddit | Jev Playing Pokemon Red Live (Open Source) singularity | supportingthedogs | 0 | 0 |
| 🟠 reddit | Kev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end LocalLLaMA | Top-Evidence174 | 34 | 7 |
| 🟠 reddit | Mica v0.1 4B got an iron pickaxe in real Minecraft without generating a single token LocalLLaMA | Top-Evidence174 | 274 | 43 |
| 🟠 reddit | Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time LocalLLaMA | Top-Evidence174 | 24 | 6 |
| 🟧 hn | Jev4j | jsumrall | 1 | 1 |
| 🟠 reddit | CLM-v0.1-8B ported to MLX — frozen Qwen3-8B encoder for instant on-device decisions, 99% top-1 agreement with the original vLLM server LocalLLaMA | WebAssemblyMan | 0 | 1 |
| 🟧 hn | Typed-lm: a Rust jev open source alternative | andrelgcclaudin | 3 | 1 |
| 🟠 reddit | An open-source alternative to Jev artificial | IceBergRock | 54 | 6 |
| 🟠 reddit | Is Jev worth the hype? singularity | beasthunterr69 | 38 | 57 |
| 🟧 hn | Show HN: Ephemeral runner for JEV-style models | zatsepin | 1 | 0 |
| 🟠 reddit | Jev already has an open-weight competitor - Deem 9b singularity | Pokenhagen | 83 | 20 |
| 🟧 hn | Turning GLM-5.3-Flash into a Jev-like decision model | flxflx | 129 | 58 |
| 🟧 hn | Von, an Open-Source Jev Alternative | k__ | 2 | 0 |
| 🟧 hn | Show HN: A local alternative to Jev – 94% on Banking77 | nico | 1 | 0 |
| 🟧 hn | Credence – Jev-style typed decisions from a local GGUF model | naughtiusmax | 1 | 0 |
| 🟠 reddit | GLM-5.3-Flash works as a Jev-like decision model with the same accuracy and speed singularity | yogthos | 86 | 16 |
| 🟠 reddit | SupersonicLabs/Julia-1 · Hugging Face LocalLLaMA | ThePrimeClock | 61 | 32 |
| 🟠 reddit | Unofficial Jev plugin for coding agents: best practices, an API reference, and 150+ community projects. Evals included. artificial | aaddrick | 1 | 2 |
| 🟧 hn | Eikos - OSS Jev-like model | emersonrsantos | 1 | 0 |
| 🟠 reddit | Should you let JEV make all of your decisions? LocalLLaMA | tenkei_01 | 0 | 26 |
| 🟧 hn | Show HN: Matching Jev on BANKING77 at a thousandth of the cost | KasianFranks | 3 | 0 |
| 🟠 reddit | 10 Technical Questions About Jev LocalLLaMA | Prashant-Lakhera | 0 | 2 |
| 🟠 reddit | I built a CLI for Jev-style typed decisions that can also run with local models LocalLLaMA | muthuishere2101 | 0 | 0 |
| 🟠 reddit | Laya: replace LLM-as-a-judge with a 322M-parameter decision engine (26,639 stars in 9 days, hands-on test) LocalLLaMA | AIFrontierReads | 0 | 4 |
| 🟠 reddit | Qwen company already rushed out a Jev competitor. No open weights yet. LocalLLaMA | pneuny | 87 | 60 |
| 🟧 hn | How to build a Jev-style classifier with DiffusionGemma and vLLM | teleforce | 1 | 1 |
| 🟧 hn | Show HN: Jevpipe – a fast System 1 for AI agents, as a Unix pipe | bothlabs | 2 | 3 |
| 🟠 reddit | ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench LocalLLaMA | Educational-Care7867 | 198 | 61 |
| 🟧 hn | Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs | pragmatwice | 3 | 0 |
| 🟧 hn | Show HN: Decide – Jev decisions in the shell, scripts, and agent skills | vsekhar | 2 | 0 |
| 🟠 reddit | Gevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring LocalLLaMA | inawhole | 0 | 8 |
| 🟠 reddit | I built a Unix pipe that lets Claude Code use Jev as a fast System 1 ClaudeAI | bothlabs | 0 | 8 |
| 🟧 hn | We swapped our LLMs for Jev. It's 39% cheaper | gczh | 7 | 4 |
| 🟠 reddit | Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) LocalLLaMA | Usual_Maximum7673 | 36 | 11 |
| 🟧 hn | Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms | firelex | 572 | 217 |
| 🟧 hn | Classify 6700 pages for $1 with a Jev-compatible VLM API | fzysingularity | 2 | 0 |
| 🟠 reddit | I routed Claude Code tasks with a cheap decision model. A cost saving is not established. ClaudeAI | Monglong_korea | 1 | 1 |
| 🟧 hn | Stanford and Nvidia's open Jev-like model | oscarfr | 1 | 0 |
| 🟧 hn | Jeeves. Reasoning improves Jev-like decision models | nicowaltz | 242 | 94 |
| 🟧 hn | What TypeSafe Got Right with the Jev Launch | nkko | 2 | 0 |
| 🟧 hn | Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound | tgluck | 65 | 15 |
| 🟧 hn | Jev-compatible document classification with open-weight VLMs | fzysingularity | 1 | 0 |
| 🟠 reddit | Decisions API OpenAI | Valuestudent | 3 | 4 |
| 🟠 reddit | TypeNotSafe singularity | Technical-Will-2862 | 458 | 81 |
| 🟧 hn | OpenJev: An open-source, Jev-compatible System One decision engine | rdudekul | 2 | 0 |
| 🟧 hn | OpenAI Launches Decisions API | dvrp | 6 | 2 |
| 🟧 hn | OpenAI Answers TypeSafe's Jev with a Decision API Built on Luna | yawnxyz | 6 | 1 |
| 🟠 reddit | Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included LocalLLaMA | One_Temperature5983 | 0 | 5 |
| 🟠 reddit | I cut cost and latency on my search engine with Jev [D] MachineLearning | AccomplishedEvent273 | 1 | 0 |
| 🟧 hn | OpenAI Launches Decisions API for Fast AI ChoicesRetrieved article excerptOpen article · Retrieved 2026-09-30T05:26:16.827020+00:00 [@OpenAIDevs](https://x.com/OpenAIDevs)
[OpenAI Developers](https://x.com/OpenAIDevs)
[OpenAI](https://x.com/OpenAI)
[@OpenAIDevs](https://x.com/OpenAIDevs)
[10h](https://x.com/OpenAIDevs/status/2105003318917697873)
Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna.
Define questions and possible answers to classify content, route requests, or choose an agent’s next action.
Available in limited preview.

00:00
[170](https://x.com/OpenAIDevs/status/2105003318917697873)
333
4.6K
451K | soltanov | 2 | 4 |
| 🟧 hn | Using Jev as a Search Reranker | dustincoates | 1 | 0 |
| 🟧 hn | Show HN: Halv cut AI agent cost by 57.1% using Jev | villaspedro | 2 | 0 |
| 🟧 hn | Laya: Multilingual, non-autoregressive System 1 decision engine | handfuloflight | 1 | 0 |
| 🟧 hn | Jev, Checked | wilsonwu8 | 1 | 0 |
| 🟧 hn | Show HN: Snap – local first decision engine open weight | emnlmn | 1 | 0 |
| 🟠 reddit | Benchmarking small confidence scoring decision (Jev, Laya) models [P] MachineLearning | iam_gkrishna | 1 | 0 |
| 🟠 reddit | NIRNAY: 450M decision model beats Jev on Banking77, runs on CPU LocalLLaMA | BrilliantSecret143 | 0 | 7 |
| 🟠 reddit | new system 1.5 model beats jev by a long shot on all 4 baselines and is #1 on image jev bench (yes its multi modal) singularity | boneMechBoy69420 | 35 | 40 |
| 🟧 hn | Clef: Open-source decision models, and new RL fine-tuning platform | jasondavies | 626 | 217 |
| 🟠 reddit | new system 1.5 model beats jev by a long shot on all 4 baselines and is #1 on image jev bench (yes its multi modal) artificial | boneMechBoy69420 | 3 | 4 |
| 🟠 reddit | Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV) LocalLLaMA | Altruistic_Heat_9531 | 0 | 12 |
| 🟧 hn | Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers | tomncooper | 135 | 49 |
| 🟧 hn | How accurately calibrated is Jev? | dblack12705 | 6 | 3 |
| 🟠 reddit | Built an open-source tool that applies to internships for me ClaudeAI | q3ndi | 2 | 3 |
| 🟠 reddit | Jev: Not Frontier, But Still Worth Your Attention LocalLLaMA | enn_nafnlaus | 0 | 5 |
2026-10-04T00:28:59Z
The 16,379-request independent probe (not frontier but Pareto-useful, strong choice-ordering sensitivity, unlikely a rebranded preexisting model) confirms the settled verdict at the largest scale yet, completing the evidence picture on all three hypothesis prongs. With calibration and frontier-comparability disproved by convergent independents, the short-task cost edge established, the category absorbed by incumbent and open ecosystem, periphery frozen since Clef, and true-zero attention (0.0 pts/h — the magnitude valve is reading the case's historical footprint, not current movement), the episode closes as absorbed; the pending Decisions API pricing/calibration watch transfers to its own episode.
2026-10-04T00:24:10Z
evidence attached: reddit.post.1wx1htr — Independent 16,379-request benchmark and probing of Jev directly tests the 'frontier-comparable' claim — finds not-frontier but Pareto-useful with odd ordering sensitivity; the corroboration/contradiction the case needs.
2026-10-03T05:48:35Z
grounded: converges/high — The world has independently arrived where Scott's canon already stood and where he is already building: the die-test/jevals/Red Hat refutation of Jev's calibrat
2026-10-03T05:40:23Z
The flagged 'grassroots adoption' attachment does not survive inspection: reddit.post.1wwclkk's body describes an internship auto-applier built on Fable/Sonnet/Opus inside Claude Code with no Jev, TypeSafe, or typed-decision content, so the case gains no new adoption signal and the prior attach rationale appears to describe a different post. With attention at the floor (~0.3 pts/h, ~470h age), all remaining movement in the already-priced Clef/TypeNotSafe tails, and zero new implementation periphery, the case's meaning is unchanged: settled judgment with the incumbent-absorption watch and triggers armed.
2026-10-03T03:23:50Z
evidence attached: reddit.post.1wwclkk — A builder organically routing all fixed-set decisions in a shipped agent tool through typesafe/jev-1.13 on OpenRouter is grassroots adoption evidence for the JEV-class decision-model case.
2026-10-02T20:19:52Z
The independent calibration audit lands as yet another convergent negative on the already-refuted calibration claim rather than contesting the verdict, so the case's meaning is unchanged: settled judgment, incumbent-absorption watch, triggers armed. Attention is flat (~11 pts/h, cooling) with zero new periphery; the 82nd-percentile peer reading is the priced Clef and Decisions-API-reaction tails, not fresh spread, so the magnitude valve no longer justifies medium.
2026-10-02T19:29:57Z
evidence attached: hn.story.49934399 — Independent calibration audit of Jev directly tests the case's core claim of calibrated probabilities — the strongest kind of contrary evidence for re-judging it.
2026-10-02T14:54:21Z
Red Hat's independent benchmark adds the first name-brand institutional negative to the value-hypothesis stack — decision models don't beat LLM-as-judge or conventional classifiers — converging with the die-test/jevals/gold-label record rather than changing the verdict; the case's open meaning is now purely incumbent-absorption watch (Decisions API specifics and TypeSafe's calibration answer, both static). Third consecutive cooling pulse (56.5→26.5→10.7 pts/h) with the loud spread valve (90th peer percentile, 3 platforms, magnitude eligible) concentrated in the already-priced Clef release tail and zero new implementation periphery this cycle, so heat drops to low despite the valve — the remaining catalysts are better served by armed triggers than sustained attention.
2026-10-02T14:26:17Z
evidence attached: hn.story.49933476 — Independent Red Hat benchmark claiming Jev-class decision models don't beat LLM-as-judge or classifiers directly contests the Jev value hypothesis.
2026-10-02T05:11:22Z
Fourth consecutive thin-composition pulse, now cooling: the velocity spike is the already-priced Clef release still climbing (~3x its p90), and the only new content since the last look — a grammar-constrained-decoding argument — re-instances the known commodity-alternative category rather than adding a fact, so no belief update and no material change. The spread reading is still loud (magnitude valve, 97.7th peer percentile across three platforms) even as momentum halves (26.5 pts/h from a 56.5 peak), which holds heat at medium rather than low; the high bar remains Decisions API specifics, unfetched since the 09-30 preview.
2026-10-02T04:27:22Z
evidence attached: reddit.post.1wvjqy8 — Argues commodity grammar-constrained generation in vLLM/llama.cpp already delivers Jev-style typed classification zero-shot, materially undercutting the specialist-Jev-model premise.
2026-10-01T22:30:11Z
Second acceleration pulse in one day, composition-identical to the first: the skeptically received wity-1 vendor claim tripped the velocity sensor by cross-posting and climbing (32→39 pts) on top of the already-priced Clef release — engagement doubled (56.5 pts/h, 28 comments/h, 99.6th percentile) but no new fact arrived, so this is a hold, not a revision; heat stays medium because the periphery is adding thin replica posts, not new implementations, and the stated high-heat bar remains Decisions API specifics (pricing/calibration/image scope), still unfetched since the 09-30 preview.
2026-10-01T20:35:47Z
evidence attached: reddit.post.1wv71c9 — Thin self-published claim, but a new multimodal reasoning decision model claiming to beat Jev on its baselines directly bears on the Jev-class competitive landscape.
2026-10-01T19:44:23Z
Meaning shift is consolidation, not hypothesis revision: Clef (open-source decision models plus an RL fine-tuning platform, attributed to Cloudflare in attach) is the first major-vendor institutional entry into the open decision-model layer and speaks directly to the closed-API can't-fine-tune weakness the calibration refutation exposed, while NIRNAY (self-admitted fine-tune-vs-zero-shot Banking77 claim) and wity-1 (unverified multimodal vendor claim, skeptically received) are thin re-instances of the saturated replica/competitor categories. Breaking from two cycles of low pricing because the numbers line jumped ~100x (34.7 pts/h vs 0.33, 14 comments/h, 99th peer percentile, magnitude valve eligible, steady momentum) on that single release — that earns medium attention, not a belief update; high waits on Decisions API specifics, not on another vendor entering a category we already know is moatless.
2026-10-01T18:31:30Z
evidence attached: hn.story.49923692 — Cloudflare shipping open-source decision models plus an RL fine-tuning platform is a major-vendor entry into the Jev-class decision-model category.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv67r7 — Unverified vendor claim that a reasoning decision model beats Jev on all four baselines — competitive pressure on the Jev category, not corroboration.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv51nt — Released open 450M decision model claiming to beat Jev 1.13.0 on Banking77 is direct competitive evidence against Jev's specialist-accuracy moat.
2026-10-01T17:53:55Z
The flagged 'substantive' object resolves to a 1-point, zero-comment independent Jev-vs-Laya benchmark whose takeaway (write label descriptions from input content, not intent) is practical label-engineering color consistent with existing doctrine — it re-instances the saturated independent-benchmark category without adding meaning: still 'incumbent confirmed, specifics pending.' Magnitude-valve eligibility again reads as accumulated footprint (277 objects/3 platforms) rather than expanding periphery: velocity is ~0.33 pts/h against a ~3054 peak with zero comment flow, so heat stays low and escalation rides the unlanded Decisions API pricing/calibration/image specifics (trigger 2), not standing attention.
2026-10-01T16:32:01Z
evidence attached: reddit.post.1wu3wep — Independent benchmark of hosted Jev against an open alternative (Laya), with concrete label-design findings, is needed context when re-judging Jev-class structured decisions.
2026-10-01T15:22:50Z
Two thin additions — an independent 'Jev, Checked' implementation article and one more open-weight System One replica (Snap) — thicken ledgers that are already saturated without adding meaning: still 'incumbent confirmed, specifics pending.' Aggregate engagement sits in terminal decay (~0.5 pts/h vs ~3060 peak, zero comments/h) and the periphery inflow is single-point re-instances of established categories, so the magnitude valve again reads as accumulated footprint (276 objects/3 platforms), not expanding spread; heat stays low and escalation rides the unlanded Decisions API pricing/calibration/image specifics (trigger 2), not standing attention.
2026-10-01T14:32:27Z
evidence attached: hn.story.49921892 — Independent released local open-weight drop-in replacement for the TypeSafe System One API — third-party ecosystem evidence that Jev-class single-pass typed decisions are spreading beyond TypeSafe itself.
2026-10-01T14:32:26Z
evidence attached: hn.story.49921076 — Independent third-party article implementing a Jev decision layer is early engagement evidence for the Jev-class structured-decisions episode.
2026-10-01T09:36:26Z
The sensor's substantive flag resolves to a 1-point, zero-comment HN re-submission of Laya — a project already central to this record — so the 'is this a class?' question it gestures at was settled days ago by the 25+ alternative catalog; meaning is unchanged ('incumbent confirmed, specifics pending'). Magnitude-valve eligibility again reads as accumulated footprint (274 objects/3 platforms) rather than expanding periphery: velocity is ~0.17 pts/h against a ~3130 peak with zero comments/h, so heat stays low and escalation rides the unlanded Decisions API pricing/calibration/image specifics (trigger 2), not standing attention.
2026-10-01T09:23:35Z
evidence attached: hn.story.49919517 — An independently released open-source non-autoregressive structured-decision engine is material context for whether fast Jev-style decision models are becoming a class rather than one vendor's claim.
2026-09-30T23:33:52Z
The Halv 57.1% agent-cost-cut is a second independently measured production saving of the same class as the earlier 39% swap — the adoption ledger thickens, the meaning doesn't: still 'incumbent confirmed, specifics pending.' The velocity spike is a p90 crossing by the plateauing TypeNotSafe thread while the aggregate line keeps cooling (~5.7 pts/h, ~0.2% of peak, comments near zero), so the magnitude valve again reads as 17 days of accumulated footprint rather than an expanding periphery — heat stays low and escalation rides the unlanded Decisions API pricing/calibration trigger, not standing attention.
2026-09-30T21:37:57Z
evidence attached: hn.story.49913080 — Third-party (Halv) reports a measured 57.1% agent cost cut using Jev on SWE-rebench — adoption evidence for the significant Jev case.
2026-09-30T20:18:02Z
Meaning holds at 'incumbent confirmed, specifics pending': the TypeNotSafe thread's 282→381 climb is single-thread amplification of the already first-party-confirmed Decisions API announcement, while Elastic Search Labs' production ecommerce reranker adoption extends the independent-adoption ledger from indie swaps to a named search vendor — a new implementation fact for the record, but it re-confirms the established short-task cost/latency edge rather than changing any conclusion. Heat stays low despite the magnitude valve: the multi-platform span is two-plus weeks of accumulated footprint (272 objects), current velocity is ~7.5 pts/h (~0.2% of peak), and the 93.5th-percentile reading is one consolidating thread riding the incumbent wave, not an expanding periphery — escalation rides the unlanded Decisions API specifics via triggers, not standing attention.
2026-09-30T18:42:39Z
evidence attached: hn.story.49911330 — Elastic Search Labs adopting Jev as an ecommerce reranker is external vendor adoption of Jev-class typed decisions for a real workload, beyond TypeSafe's own claims.
2026-09-30T05:55:27Z
First-party receipt landed for the incumbent sub-episode: OpenAI's own Devs announcement confirms Decisions API on GPT-6 Luna in limited preview with Jev's exact typed question/answer shape (451K-view tweet), upgrading 'community-framed answer' to 'official second supplier' — but pricing, calibration approach, and image-input scope remain unverified, so nothing escalates. Heat drops to low: velocity is sub-1% of peak and cooling, the ~90th-percentile magnitude reading is 16 days of accumulated footprint rather than an expanding periphery (one consolidating thread creeping 267→282, thin 2-6 pt re-posts), and the remaining follow-up (Decisions API pricing/calibration specifics, TypeSafe's calibration answer) is carried by triggers, not standing attention.
2026-09-30T05:42:54Z
evidence attached: hn.story.49904454 — OpenAI entering typed-decision serving with a GPT-6 Luna-backed API is direct competitive pressure on Jev's latency/cost claim and belongs on that significant case, without displacing TypeSafe's own artifact as anchor.
2026-09-30T04:26:46Z
grounded: converges/high — Scott's most production-adjacent radar case: his dev wiki already carries a Jev warm-start (dev:project.jev — measured doctrine, jevkit, two harvested productio
2026-09-30T04:18:43Z
Latest deltas are amplification, not substance: the TypeNotSafe incumbent thread crept from 264 to 267 points and a second thin adoption anecdote (DeepSeek-extraction swap via OpenRouter) re-confirms an already-established cost/latency record rather than extending it. Meaning holds at 'incumbent answered, specifics pending' — the 88th-percentile spread reading reflects 16 days of accumulated footprint, not current velocity (~0.8% of peak, cooling), so medium heat now prices solely the unlanded Decisions API details, and a next quiet look should drop this to low.
2026-09-30T03:31:23Z
evidence attached: reddit.post.1wtsn3s — First independent adoption report: a builder swapped DeepSeek extraction for TypeSafe Jev via OpenRouter for cost/latency, directly supporting the Jev case.
2026-09-30T02:20:18Z
Since the Decisions API trigger fired, no load-bearing details have landed: the incumbent sub-episode is consolidating into a single community thread ('TypeNotSafe', 220 pts/61 comments) while the fresh deltas are amplification (ImaJev-4b velocity spike on an already-recorded noisy topper) and a further open vision-capable replication (typevet) in an already vision-saturated catalog. The case's meaning holds at 'incumbent answered, specifics pending' — it neither advances toward escalation (no pricing/calibration/latency specifics from OpenAI) nor decays to low (the sub-episode object still runs top-decile for its cohort), so significant/medium stands on follow-up value rather than aggregate velocity, which is at ~1.4% of peak and cooling.
2026-09-30T00:38:31Z
evidence attached: reddit.post.1wtmz2p — An open vision-capable replication of Jev-style typed decisions with per-label probabilities on a 4090 directly contextualizes TypeSafe's claimed differentiation and extends beyond its text-only scope.
2026-09-29T22:59:43Z
The case's standing frontier-lab trigger fired: OpenAI shipped a Luna-based Decisions API at DevDay, independently framed as its answer to Jev — typed decisions are now an incumbent-native product category, capping Jev's differentiation from the top just as the 25+ open-alternative catalog caps it from below. The case's forward meaning shifts from adjudicating Jev's claims (settled: cost/latency edge real, calibration refuted) to tracking incumbent consolidation of the category, with Jev's residual edge narrowed to short-request cost/latency pending Decisions API pricing and calibration details.
2026-09-29T20:53:12Z
evidence attached: hn.story.49896979 — Independent coverage explicitly framing the Decisions API as OpenAI's answer to Jev — key competitive evidence for the open Jev case.
2026-09-29T20:53:12Z
evidence attached: hn.story.49897587 — OpenAI's Decisions API launch is the incumbent artifact of the Luna-based counter-move and directly contests Jev's cheap-structured-decision claim.
2026-09-29T20:53:12Z
evidence attached: hn.story.49898615 — An open-source Jev-compatible decision engine appearing right after TypeSafe's announcement is spread evidence the Jev pattern is being replicated outside one vendor.
2026-09-29T20:53:11Z
evidence attached: reddit.post.1wti0rp — Credibility backlash post whose top comment reports OpenAI shipped a competing Decisions API on Luna — material competitive and validation context for the TypeSafe Jev case.
2026-09-29T20:53:11Z
evidence attached: reddit.post.1wticxm — Thin question thread but the batch's only attestation that OpenAI surfaced a 'Decisions API' at DevDay — a big-lab entry into the Jev-style structured-decision category the case watches.
2026-09-29T19:14:14Z
Meaning unchanged: the sensor's 34x velocity spike is Jevstiller moving 5.7 pts/h off a 0.17 pts/h baseline — small-object noise, not case-wide expansion — and the only new evidence is a 1-pt open-weight-VLM document-classification pipeline (same author/pattern as the earlier 6700-pages-for-$1 post), one more entry in the already-saturated Jev-compatible catalog; Jeeves keeps accumulating engagement (559 pts) on the adoption line priced at the last two looks. Consolidation decay holds (~32 pts/h vs ~3503 peak across 263 items at 388h), so heat stays low despite the magnitude-valve flag, which here measures historical launch/replication breadth on two saturated communities, not fresh periphery — no new implementations, communities, or outlets appeared this window.
2026-09-29T17:42:00Z
evidence attached: hn.story.49896315 — A third-party Jev-compatible document-classification pipeline is early ecosystem-adoption evidence for the structured-decision pattern beyond TypeSafe's own claim.
2026-09-29T15:06:26Z
Meaning unchanged: the measured-rate uptick (23→34 pts/h) is PostHog's Jeeves accumulating engagement on the adoption line priced last look, not fresh expansion, and Jevstiller (distillation with disagreement bounds, 10 pts) is one more entry in the already-saturated catalog — no trigger fired and no established claim moved. The magnitude-valve flag and 95th-percentile peer reading reflect historical launch/replication breadth across 262 items on a cooling curve, not new cross-platform spread, so heat stays low.
2026-09-29T14:26:57Z
evidence attached: hn.story.49891769 — Third-party tooling distilling Jev into local models with disagreement bounds is early ecosystem evidence bearing directly on whether Jev becomes a real structured-decision platform.
2026-09-29T11:59:15Z
Consolidation decay continues on schedule (27.5→23.3 pts/h, cooling; the 94th-percentile peer reading reflects still-accumulating legacy items, and the magnitude-valve spread remains historical breadth — launch wave, replication wave, fatigue meta — not fresh expansion). The window's only new content is PostHog building 'Jeeves' on Jev-like decision models — the first named product company adopting the pattern internally, a modest but genuine strengthening of the already-established adoption line at 29 pts — plus reception meta at 2 pts. No trigger fired and no established claim moved: this is spread evidence inside an adjudicated case, not a repricing event. Heat stays low; the reheating paths remain TypeSafe's rebuttal, a frontier-lab native endpoint (Qwen decision-model-preview, still docs-only), or a matched head-to-head on Scott's tasks.
2026-09-29T11:27:12Z
evidence attached: hn.story.49891067 — Substantive third-party analysis of the Jev launch itself; reception context for an open significant case, though not independent corroboration of its performance claims.
2026-09-29T11:27:11Z
evidence attached: hn.story.49891290 — PostHog building 'Jeeves' on Jev-like decision models is independent third-party engagement with the Jev pattern, relevant spread evidence for the significant TypeSafe case.
2026-09-29T10:58:00Z
Consolidation decay confirmed, not a new chapter: the Stanford/NVIDIA Jev-like entry is provenance-notable but traction-nil (1 pt, 0 comments) and mechanism-adjacent (action caching, not typed decisions) — one more name in an already-saturated catalog, and the window's only velocity spike is the already-flagged ImaJev-4b self-reported topper. With the rate halved again (62→27.5 pts/h), momentum cooling, and all three triggers unfired, the case's meaning is settled enough that attention — not belief — should step down: heat to low; TypeSafe's answer, a frontier-lab native endpoint, or a matched head-to-head are the only reheating paths.
2026-09-29T10:24:30Z
evidence attached: hn.story.49890520 — An open Stanford/NVIDIA Jev-like model caching reusable agent actions at up to 9x speed directly changes the competitive context for whether Jev-style structured decisions become a category.
2026-09-29T04:43:58Z
First validated live third-party routing gate in Claude Code (jev-gate) ships with an honest null — no cost saving demonstrated in its only tested cell (every task landed the same tier) — giving the flagship model-routing use case its first end-to-end counter-datapoint and sharpening the measure-cost-per-accepted-decision doctrine, without moving any core claim. Heat holds medium against a 99.8-percentile 'accelerating' speedometer because the rate (~62 pts/h, unchanged from the last look) decomposes to the already-priced Jeff/Ollaya/Jev-in-25-Lines wave tail plus astroturf-fatigue meta rather than fresh expansion; the only new substantive item this window was a 1-point post, and all three triggers remain unfired — medium keeps the frontier-lab trigger (Qwen decision-model-preview, still docs-only) under active watch without re-alerting on known breadth.
2026-09-29T04:26:18Z
evidence attached: reddit.post.1wt086x — First observed third-party artifact shipping Jev as a live routing gate in Claude Code, with the author honestly reporting no cost saving demonstrated — early adoption and counter-evidence the case's re-judge should weigh.
2026-09-29T03:31:04Z
No interpretive change: the 6700-pages-for-$1 post is one more third-party cost confirmation in an already-ruled saturated catalog — the per-request cost edge was already independently established (production swap report, Tessl/Near Here benchmarks). The 62 pts/h re-acceleration and 98.8 peer-percentile decompose to the Jeff HN post still climbing (162→327) atop the existing periphery; momentum is cooling and the community signal remains astroturf-fatigue-tinged, so magnitude-valve eligibility reads as historical breadth, not fresh expansion. Heat holds at medium: live motion justifies continued attention on the stirring frontier-lab trigger (Qwen decision-model-preview 89/59, still docs-only; TypeSafe still publicly unanswered on calibration), but no claim moved and all three triggers are unfired.
2026-09-29T02:30:36Z
evidence attached: hn.story.49887163 — Third-party adoption evidence — 6700 pages classified for $1 via a Jev-compatible API supports the cost side of the Jev hypothesis beyond TypeSafe's own claims.
2026-09-28T23:42:58Z
No interpretive change: the only new evidence is the HN front-page version (162 pts) of the already-adjudicated Jeff 0.8B/2B home-trainable models — a further no-accuracy-moat confirmation adding breadth, not meaning. But the re-acceleration to 52.5 pts/h now decomposes across fresh multi-platform periphery (Jeff HN 162, sustained Ollaya 578, Mica Minecraft 274, ImaJev 141, Qwen decision-model-preview discussion 87/58) rather than the flagged ImaJev post alone, so heat rises low→medium to keep attention on the stirring frontier-lab trigger (still docs-only); no claim moved and all three triggers remain unfired.
2026-09-28T23:28:17Z
evidence attached: hn.story.49883844 — Independent home-trainable Jev-compatible 0.8B decision model at ~30ms corroborates the small-decision-model-as-routing-component thesis from outside TypeSafe.
2026-09-28T22:46:38Z
No interpretive change: the locally-trained Jeff 0.8B/2B one-pass models are the ~22nd entry in an already-ruled saturated replication catalog, confirming no accuracy moat at the small end; the measured 'acceleration' (24 vs ~19 pts/h) and 90.8 peer-percentile are base-rate artifacts off a decayed base, driven by the karma-solicited ImaJev post accruing points and a 9-point catalog post. Heat stays low despite magnitude-valve eligibility — the multi-platform spread is historical (255 objects over 15 days), and the current community signal is astroturf fatigue (1182-pt 'stop the Jev ads' meta-post), not expansion. All three triggers unfired (Qwen decision-model-preview still docs-only).
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wspn24 — Locally trained open-weight 0.8B/2B one-pass decision models matching Jev's published score at ~30ms directly bear on Jev's differentiation and the viability of the category.
2026-09-28T20:21:53Z
The cost thesis gained its first independent production corroboration — a documented LLM→Jev swap saving 39% (author-confirmed on HN) — upgrading bounded cost utility from benchmark-observed to production-observed; everything else this window is catalog churn (Gevva0's self-reported JevBench #1 via logit scoring with unpublished data, a duplicate jevpipe post) plus a velocity spike that is just the astroturf-flagged ImaJev post accruing points. The adjudication is otherwise unchanged and all triggers remain unfired (Qwen decision-model-preview still docs-only).
2026-09-28T18:36:54Z
evidence attached: hn.story.49881537 — Independent third-party production report of 39% savings swapping LLMs for Jev is the first outside corroboration of the case's cost claims.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wsiu1y — Third-party tooling plus a cited independent Jev-vs-LLM benchmark is adoption and calibration evidence for the Jev case.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wshwqz — Independent community build of a JEV-like calibrated decision engine on open Gemma topping JevBench at ~214ms — direct replication-context for whether the structured-decision pattern generalizes beyond TypeSafe's model.
2026-09-28T17:22:12Z
Twentieth look, meaning unchanged and dormancy re-confirmed: 3.5 pts/h against a 3352 peak at ~363h age, and the peer-percentile rebound to 75.0 is a stale-cohort artifact, not renewed spread. The new items are catalog churn — ImaJev-4b's self-reported JevBench #1 (datasets unpublished, from an account that earlier solicited karma to post it) adds no measured signal, while Jeva.cpp and Decide are zero-traction entries in the already-catalogued interface-standardization trend. The case remains an armed low-heat watch over a settled adjudication; trigger 2 (frontier-lab typed-decision endpoints) still stirring.
2026-09-28T15:45:04Z
evidence attached: hn.story.49878492 — Independent corroboration: third-party CLI tooling built on TypeSafe's jev-latest for shell and agent-skill decisions is direct adoption evidence for Jev decision models.
2026-09-28T15:45:04Z
evidence attached: hn.story.49878572 — Released llama.cpp fork exposing a JEV-compatible API for any local model is ecosystem-spread evidence for whether Jev decision interfaces become a standard layer.
2026-09-28T15:45:04Z
evidence attached: reddit.post.1wsgrma — Independent community 4B decision model topping JevBench and beating GPT-5.6 Luna on DecisionBench (self-reported) is direct support for the frontier-comparable-cheap-decision-model thesis.
2026-09-28T14:15:49Z
Nineteenth look, meaning unchanged and dormancy confirmed: absolute velocity is 0.17 pts/h vs a 3313 peak at ~360h age and peer percentile fell from 92.9 to 34.3, validating the prior stale-cohort-artifact discount of the magnitude-valve reading; the only new evidence is Jevpipe (2 pts/2 comments), another low-traction wrapper in the saturated catalog, and Qwen-preview comment churn advances trigger 2 not at all. The case consolidates as an armed low-heat watch over a settled adjudication, with trigger 2 (frontier-lab typed-decision endpoints) still stirring.
2026-09-28T13:35:16Z
evidence attached: hn.story.49877434 — Independent 'System 1' Unix-pipe tool is early, weak ecosystem-spread evidence for the Jev structured-decision paradigm.
2026-09-28T06:45:06Z
Eighteenth look: the review screen's 'material development (noul=0.89)' flag is a false positive on a probability figure inside the Laya echo post's updated discussion — the actual delta is a third-party correction that Laya's ~33ms GPU-headline latency runs ~3.1s/forward on a 2-vCPU CPU box (~100× slower), a catalog footnote tempering one of ~19 alternatives, not a case-level fact. Adjudication and all three triggers unchanged; low heat stands despite the 92.9th-percentile/magnitude-valve readings because absolute velocity is 11 pts/h vs a 3307 peak at 353h age and the periphery yields only repeats and footnotes, not new implementations, communities, or outlets — the percentile is a stale-cohort artifact, not spread.
2026-09-28T03:40:44Z
Seventeenth look, meaning unchanged: the Qwen decision-model-preview surface is drawing sustained discussion (72 pts/45 comments, 17× peer-baseline velocity) — community amplification of trigger-2's stirring, not trigger advancement (still no announcement, open weights, or pricing); the only other delta is a zero-traction DiffusionGemma+vLLM tutorial, another entry in the saturated catalog. Adjudication unchanged; case stays armed at low heat.
2026-09-28T03:25:10Z
evidence attached: hn.story.49872836 — Third-party tutorial reproducing Jev-style structured-decision classifiers on open tooling (DiffusionGemma + vLLM) is early ecosystem-spread evidence contextualizing whether the pattern generalizes beyond TypeSafe's claim.
2026-09-28T00:39:48Z
Sixteenth look: the standing 'frontier labs ship native typed-decision endpoints' trigger stirred for the first time — Qwen quietly published a decision-model-preview docs page (no announcement, open weights, or pricing yet), converting that strategic risk from pure analysis to an early product surface. Core adjudication unchanged (bounded utility, calibration refuted, parity unproven); the Laya 26k-star post is a zero-traction, unverified diffusion echo, so heat stays low while this first concrete trigger movement earns the material flag and keeps the trigger-2 watch alive.
2026-09-28T00:23:46Z
evidence attached: reddit.post.1wrz9py — Qwen quietly shipping a Jev-style decision model is competitive-response evidence that major labs treat typed structured decisions as a real product category.
2026-09-28T00:23:46Z
evidence attached: reddit.post.1wrzjib — A 26k-star open 322M typed-decision engine with hands-on testing is independent parallel evidence that the typed-decision category Jev represents is real and spreading.
2026-09-27T20:58:59Z
Fifteenth consolidation look, meaning unchanged: the only delta is a zero-traction jevx CLI post (reddit.post.1wrs4i5) — another wrapper entry in the already-saturated open-alternative catalog, not a new fact, result, or contradiction. Velocity continues its decay tail (~6.7 pts/h vs ~3272 peak at ~343h age, cooling); the 72.7th-percentile peer reading and magnitude-valve flag remain stale-cohort artifacts of the accumulated 244-object corpus, and r/LocalLLaMA fatigue/astroturf complaints persist. All three standing triggers unfired, so the case stays armed at low heat.
2026-09-27T20:29:18Z
evidence attached: reddit.post.1wrs4i5 — Third-party CLI adopting Jev-style typed decisions with local models is early ecosystem-diffusion evidence for the open case, though the post itself has zero traction.
2026-09-27T16:38:03Z
Fourteenth consolidation look, meaning unchanged: the new '10 Technical Questions About Jev' post (reddit.post.1wrn9zf) restates open questions the case already catalogs — undisclosed architecture, RLCD-likely-not-RL, TypeSafe's public silence on the calibration refutation — and arrived with zero traction (0.08 ratio), adding no fact or shift. Measured heat now correctly reads cooling (12.5 pts/h vs 3272 peak at ~338h), confirming the decay-tail diagnosis; the magnitude-valve flag remains an artifact of the accumulated 243-object launch corpus. All three standing triggers unfired.
2026-09-27T16:26:17Z
evidence attached: reddit.post.1wrn9zf — Technical scrutiny of Jev's undisclosed architecture and training directly contextualises how much weight TypeSafe's frontier-comparable claims can carry.
2026-09-27T15:27:36Z
Thirteenth consolidation look, meaning unchanged: the Tuatara embedding-model result (91.79% vs Jev's 92.40% on Jev's own 3,080 Banking77 test messages, claimed 1/1000 cost) is a second conventional-classifier statistical tie on the same public benchmark — the first receipt aimed at the cost claim rather than accuracy, but it reinforces the already-catalogued finding (embedding+logreg 94.25% Banking77) rather than shifting any doctrine or trigger. Heat stays low despite the magnitude valve: that spread reading is the accumulated 242-object launch corpus within HN/Reddit only, current velocity is a ~17 pts/h decay tail at ~337h age, the 'accelerating' flag and 87.5th-percentile peer reading are stale-cohort artifacts, and r/LocalLLaMA is in open fatigue/mod-complaint mode — all three standing triggers unfired.
2026-09-27T15:24:21Z
evidence attached: hn.story.49867086 — Independent third-party benchmark claims matching Jev on BANKING77 at a thousandth of the cost, direct countervailing evidence on Jev's cost/latency value proposition.
2026-09-27T14:37:59Z
Twelfth consolidation look, meaning unchanged: the new independent gold-label benchmark of Jev vs frontier models on bounded-decision contracts (reddit.post.1wrktrp) confirms the already-established finding that Jev's accuracy is bounded and task-dependent, not frontier-comparable — but it arrived dead on arrival (0.16 upvote ratio, downvote-on-sight replies) amid visible community Jev-fatigue, so it adds a weak receipt, not a shift. The velocity-spike flag is a floor-baseline artifact (5.8× on ~1 pt/h base) from a small Julia-1 repost; all three standing triggers remain unfired.
2026-09-27T14:25:55Z
evidence attached: reddit.post.1wrktrp — Independent third-party benchmark directly testing JEV against frontier models on bounded-decision contracts, finding task-dependent gaps that qualify the 'frontier-comparable' claim.
2026-09-27T08:40:23Z
Eleventh consolidation look, meaning unchanged: the only sensor activity is marginal engagement on known evidence (the Mica Minecraft demo creeping 266→269 pts, a comment update on the GLM-5.3-Flash-as-decider story) — the velocity-spike flag is that stale demo post's cohort standing, not new spread. No new evidence arrived, all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab native typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired, and with material_change false the case should be allowed to sleep longer.
2026-09-27T04:33:38Z
Tenth consolidation look, meaning unchanged: the new evidence is a community Jev plugin ecosystem indexing 150+ projects with evals, and Eikos, another OSS Jev-like model — adoption consolidation and a further catalog entry that reinforce the established findings (broad but bounded utility, trivially replicable category, calibration refuted) without shifting them. Heat stays low despite the magnitude valve and 93rd-percentile peer reading: those remain the Mica demo's within-cohort standing and accumulated launch-corpus spread, not current velocity (~20 pts/h vs ~3267 peak at ~326h, cooling); the periphery still repeats only within HN/r/LocalLLaMA, and all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab native typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-27T04:23:10Z
evidence attached: hn.story.49863193 — An OSS Jev-like model bears on whether TypeSafe's structured-decision approach is replicable outside its closed early access.
2026-09-27T04:23:10Z
evidence attached: reddit.post.1wra3v2 — A community Jev plugin ecosystem (best practices, 150+ project index, cross-provider-judged evals) is adoption context the TypeSafe Jev case should carry.
2026-09-27T03:45:15Z
Ninth consolidation look, meaning unchanged: the only new evidence is SupersonicLabs' Julia-1, a 144M mmBERT-based open decision model — the smallest catalog entry yet, deepening rather than shifting the established finding that compact typed-decision models are a trivially replicable category with no Jev accuracy moat. Heat stays low despite the magnitude valve and 92nd-percentile peer reading: those are again the Mica demo's within-cohort standing and a launch-corpus artifact, not current velocity (~20 pts/h vs ~3262 peak, cooling); the periphery still repeats inside HN/r/LocalLLaMA and each addition reinforces an already-established conclusion, and all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab native typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-27T03:23:54Z
evidence attached: reddit.post.1wr8d4d — An explicitly Jev-like open 144M-parameter CPU decision model is independent context on whether compact structured-decision models are a replicable category or TypeSafe-specific hype.
2026-09-27T02:36:11Z
grounded: converges/high — The world has independently arrived where Scott's canon already stands: the fair-die and jevals ground-truth refutations of Jev's calibrated-confidence claim ar
2026-09-27T02:27:11Z
Eighth consecutive consolidation look, meaning unchanged: the new evidence is a low-traction r/LocalLLaMA cross-post (yogthos) showing GLM-5.3-Flash works as a Jev-like decision model — a line already cataloged via the 31-pt HN thread on the same approach — reinforcing the established 'conventional models suffice on Jev's bounded tasks' finding rather than shifting it. Heat stays low despite the magnitude valve and 92nd-percentile peer reading: that percentile is the Mica v0.1 4B Minecraft demo (~259 pts) standing within its own age cohort, not case velocity (~17 pts/h, ~0.5% of the ~3254 launch peak, momentum cooling), the periphery still repeats only inside HN/r/LocalLLaMA, and all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab native typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-27T02:22:36Z
evidence attached: reddit.post.1wr83hi — shared external link with case evidence
2026-09-27T00:55:01Z
Seventh consecutive consolidation look, meaning unchanged: the only movement is a Mica v0.1 4B Minecraft botting demo still climbing (~253 pts, the single warm object behind the velocity flag) and a one-point local GGUF implementation (Credence) joining the saturated open catalog — both reinforce the established open-alternative and real-time-use lines rather than shift them. Heat stays low because aggregate velocity halved again (~18.5 vs ~43 pts/h, under 1% of the ~3217 launch peak, momentum cooling) and all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-27T00:25:45Z
evidence attached: hn.story.49861867 — Independent local GGUF implementation of Jev-style typed decisions is early adoption evidence bearing on whether the Jev pattern spreads beyond TypeSafe.
2026-09-26T18:58:34Z
Sixth consecutive consolidation look, meaning unchanged: the Von HN entry is a duplicate of known coverage, and the one new substance — an embedding+logistic-regression pipeline hitting ~94% on Banking77 — concretely confirms the already-established line that conventional classifiers have no accuracy disadvantage on Jev's bounded tasks. That is reinforcement, not a shift, so no material change and no clock reset. Heat stays low despite the magnitude valve and a slight rate uptick (~43 vs ~37 pts/h) because velocity is ~1.3% of peak, momentum is cooling, the periphery still repeats only inside HN/r/LocalLLaMA, and all three standing triggers (TypeSafe calibration rebuttal with data, frontier-lab typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-26T18:25:02Z
evidence attached: hn.story.49858795 — A 642KB embedding+logistic-regression classifier beats Jev zero-shot on Banking77, directly bearing on whether Jev's frontier-comparable structured-decision claim has a defensible moat.
2026-09-26T18:25:02Z
evidence attached: hn.story.49858763 — shared external link with case evidence
2026-09-26T16:41:35Z
Fifth consecutive consolidation look, meaning unchanged: LibertAI's Deem 9b/0.8B (68.9% vs Jev's 74.1% on the hard benchmark) and a GLM-5.3-Flash single-forward-pass adaptation join the saturated open ecosystem without moving the established finding — open alternatives stay credible, uneven, and sub-parity. The jevos velocity spike is peer-relative noise on a 13-point object. Heat stays low despite the magnitude-valve flag because the top-decile cross-platform spread is launch-corpus history, not current velocity (~37 pts/h vs ~3200 peak, cooling, periphery repeating inside HN/r/LocalLLaMA only), and all three standing triggers (TypeSafe calibration rebuttal, frontier-lab typed-decision endpoint, matched head-to-head on Scott's tasks) remain unfired.
2026-09-26T16:26:09Z
evidence attached: hn.story.49857656 — Third-party replication of the Jev-like decision-model pattern on GLM-5.3-Flash shows the approach spreading beyond TypeSafe's release.
2026-09-26T16:26:09Z
evidence attached: reddit.post.1wqtjd4 — An open-weight Jev competitor (Deem 9b/0.8B on Qwen3.5) is material competitive context for whether Jev-style decision models become a category.
2026-09-26T15:40:57Z
Fourth consecutive consolidation look with no change in meaning: jevos is one more independent bounded measurement (Jev 0.927 on 2000 held-out rule yes/no vs jevos 0.815 on CPU and Laya 0.489 — confirming Jev's narrow accuracy is real while showing open-alternative quality is uneven and task-bound), alongside a practitioner-sentiment thread and one CLI tooling instance inside the same two communities. Decay continues on schedule (~35 pts/h vs ~3190 peak at ~314h age, cooling); the magnitude-valve and 95.8th-percentile readings remain launch-corpus/single-hot-object artifacts, periphery is not expanding, and no standing trigger (TypeSafe calibration rebuttal, frontier-lab typed-decision endpoint, matched head-to-head on Scott's tasks) has fired.
2026-09-26T15:24:09Z
evidence attached: hn.story.49857297 — First-party CLI runner for JEV-style structured-decision models with local backends is early ecosystem tooling contextualising Jev's practicality beyond TypeSafe.
2026-09-26T15:24:09Z
evidence attached: reddit.post.1wqs5d9 — Independent community thread questioning Jev hype, with a practitioner reporting real LLM-cost-cutting use — third-party adoption sentiment for the significant Jev case.
2026-09-26T15:24:09Z
evidence attached: reddit.post.1wqsawg — An open-source one-pass classifier measuring Jev at 0.927 on held-out yes/no rules and reaching 0.815 on CPU is independent validation context and competition for the Jev case.
2026-09-26T13:34:41Z
No new meaning: the substantive_evidence flag resolves to Typed-lm, a 3-point/1-comment hobby Rust alternative — a third consecutive instance of the already-recorded open-ecosystem saturation (after Jev4j and the CLM MLX port), whose sole comment restates established doctrine (Jev's edge matters only at very high request rates where accuracy can be traded; modern hosted LLMs are cheap for everything else). Decay continues on schedule (~34 pts/h vs ~3146 peak at ~311h age, momentum cooling); the magnitude-valve and 95.5th-percentile readings are launch-corpus/single-hot-object artifacts — the periphery repeats inside HN and r/LocalLLaMA rather than expanding into new communities or consequential participants, and no standing trigger (TypeSafe calibration rebuttal, frontier-lab typed-decision endpoint, matched head-to-head on Scott's tasks) has fired. Doctrine, triggers, and Scott-facing assessment unchanged.
2026-09-26T13:26:30Z
evidence attached: hn.story.49855980 — A released open-source Rust alternative to the closed early-access Jev model directly bears on whether TypeSafe's structured-decision approach gets commoditized.
2026-09-26T12:34:55Z
No new meaning: the velocity spike decomposes into same-creator Mica promotion (Minecraft/Tetris demos, with a subreddit mod flagging the creator over the self-promotion ratio) plus a CLM port to MLX of an already-recorded open model at 99% top-1 agreement — runtime porting and demo amplification, not new evidence, new communities, or new technique territory. Consolidation decay confirmed (~35 pts/h, halved again from ~42, at ~310h age); doctrine, standing triggers and Scott-facing assessment unchanged.
2026-09-26T12:23:14Z
evidence attached: reddit.post.1wqokwd — Open-weights Apache-2.0 CLM port replicates Jev's typed-decision-with-probability pattern on-device with 99% agreement, independent community evidence the approach is real and commoditizable.
2026-09-26T08:23:44Z
No new meaning: the flagged Jev4j item is a 1-point hobby Java wrapper putting Jev-style choices into switch expressions — one more instance of the already-recorded commodity-tooling saturation, not spread into new communities or technique territory; the velocity_spike resolves to comment drips on already-recorded Mica/Pokémon/Jev-vs-Kev threads. Measured velocity halved (~42 pts/h vs ~89) at ~306h age, confirming the predicted consolidation decay; magnitude-valve and 96th-percentile readings remain single-hot-object artifacts. Doctrine, standing triggers and Scott-facing assessment unchanged.
2026-09-26T08:22:35Z
evidence attached: hn.story.49854129 — Apparent independent Java implementation of Jev would be third-party ecosystem-spread evidence for the significant open case, if it targets the same technique — worth carrying with that caveat.
2026-09-25T23:46:04Z
No new meaning: the Mica v0.1 cluster ($30-trained Apache-2.0 4B, agent-loop demos, self-reported Tetris edge over Kev) is one more instance of the already-recorded near-zero-cost open replication — demos and same-creator benchmarks, not a measured result. The measured re-acceleration (~89 pts/h vs ~43, 'accelerating' flag) decomposes into Ollaya's fresh 278-point accumulation plus same-creator and Pokémon repeats inside HN/r/LocalLLaMA: consolidation decay, not spread, so heat stays low despite the velocity, 99.4th-percentile and magnitude-valve readings, which remain corpus artifacts of one hot young object plus a dense launch tail.
2026-09-25T23:26:00Z
evidence attached: reddit.post.1wqag1i — An open Apache-2.0 Jev-format decision model trained for under $30 materially contextualizes Jev's significance and moat — re-judging the case must mention it.
2026-09-25T23:26:00Z
evidence attached: reddit.post.1wqahbz — Zero-token Minecraft iron-pickaxe run demonstrates Jev-style logit decision models working in a real agent loop at ~90-150ms per decision.
2026-09-25T23:26:00Z
evidence attached: reddit.post.1wqai4b — Same-creator seeded Tetris comparison adds quantified decision-quality evidence for the open Jev-format model thread.
2026-09-25T20:29:29Z
No new meaning: the sole fresh attachment is a zero-engagement Pokémon demo repeat, and the velocity spike is the already-priced Jev-vs-Kev head-to-head re-accumulating points (~43 pts/h vs ~3054 peak, cooling); the Kev result remains the last material change. Periphery still repeats within HN/r/LocalLLaMA — the magnitude-valve flag and 96.6th-percentile peer reading stay launch-corpus artifacts — so heat remains low.
2026-09-25T20:24:27Z
evidence attached: reddit.post.1wq6l9b — shared external link with case evidence
2026-09-25T19:58:58Z
Ecosystem consolidation, not new meaning: Ollaya (an Ollama-style serving layer for open decision models) and a CLM paper cross-post are commodity-layer repeats of already-priced open-ecosystem facts inside the same HN/r/LocalLLaMA communities; the Kev head-to-head remains the last material result. Attention continues decaying (~24 pts/h vs ~3023 peak) — the magnitude-valve flag and 96.7th-percentile peer reading are artifacts of the 221-item launch corpus, not new periphery, so heat stays low.
2026-09-25T19:24:47Z
evidence attached: hn.story.49848279 — Independent research release claiming a fast, generalizable System One decision model corroborates the broader non-autoregressive structured-decision direction the Jev case tracks.
2026-09-25T19:24:46Z
evidence attached: hn.story.49848269 — A front-page Ollama-style server for open-source Jev-style decision models is third-party ecosystem/adoption evidence for the Jev paradigm the TypeSafe case tracks.
2026-09-25T18:57:47Z
The case's own Jev-vs-Kev-gaps watch trigger partially fired: a matched independent head-to-head of Jev vs open Kev 4B on 362 fresh post-launch items (published code) shows accuracy parity and Kev better calibrated — reinforcing the calibration refutation — but Kev's ~257-token per-request overhead makes short requests up to 12× pricier, so LLM-based open drop-ins pay a real token tax and Jev keeps a genuine per-request cost/latency edge on short requests. This meaningfully refines the 'increasingly satisfiable with open alternatives' line and Scott's benchmark matrix (material change), while attention remains in consolidation decay (~13 pts/h vs ~2964 peak), so heat stays low.
2026-09-25T18:28:12Z
evidence attached: hn.story.49847306 — A head-to-head test of an open-source Jev alternative is direct competitive/adoption evidence for the Jev decision-model category the TypeSafe case tracks.
2026-09-25T18:28:12Z
evidence attached: reddit.post.1wq2hfc — Independent head-to-head of Jev vs open Kev 4B with published code: accuracy parity, better calibration (supports claim), but a ~257-token per-request overhead making short requests up to 12x pricier (contradicts cost claims) — the case's first real outside evaluation.
2026-09-25T17:43:29Z
Meaning unchanged: the substantive_evidence trigger resolves to a 3-pt 'Jev Plays Pokémon Red' build — the Nth same-pattern game demo (JevEmon already walked FireRed; Pokemon-harness and 9-games runs were priced days ago) — and the engagement deltas are tail noise (CLM 406→416, one 2→3 comment bump). Consolidation decay deepens at ~5 pts/h vs ~2927 peak; the magnitude-valve and 80th-percentile peer readings remain stock artifacts of the 217-item launch corpus — the periphery is repeating inside the same HN/r/LocalLLaMA communities, not expanding into new ones, and neither load-bearing trigger (TypeSafe data-backed calibration rebuttal, frontier typed-decision endpoints) has fired.
2026-09-25T16:30:12Z
evidence attached: hn.story.49845172 — Independent open-source use of Jev in a real-time decision loop with visible token costs is weak but real external evidence for its fast structured-decision claim.
2026-09-25T14:27:52Z
Consolidation decay continues unchanged. The three new attachments are 1–2-pt same-pattern builds from the same HN communities: the escalation router and DSPy ReAnchor optimizer repeat patterns the case already prices, and the multi-model stress benchmark reinforces rather than revises the settled picture (open models competitive; calibration refuted) — independent benchmarking has become ambient background of this case, not new signal. Neither load-bearing trigger (TypeSafe data-backed rebuttal, frontier typed-decision endpoints) has fired; heat stays low on ~9 pts/h current flow vs ~2925 peak, with the 90th-percentile and magnitude-valve readings remaining stock artifacts of the 216-item launch corpus.
2026-09-25T14:25:03Z
evidence attached: hn.story.49844668 — An independent benchmark repo stressing Jev against Laya and open models is exactly the outside evaluation evidence the case needs, the strongest of today's Jev signals.
2026-09-25T14:25:03Z
evidence attached: hn.story.49844759 — Third-party tooling for building and optimizing Jev programs in DSPy contextualises whether the Jev paradigm is gaining ecosystem traction beyond TypeSafe.
2026-09-25T14:25:03Z
evidence attached: hn.story.49844757 — An independent Jev-like classifier-to-LLM escalation router is adoption evidence that the structured-decision pattern the significant Jev case tracks is spreading.
2026-09-25T13:47:30Z
Meaning unchanged. The velocity_spike resolves to engagement tail on the already-priced 'Jev isn't new tech' backlash post (750→767 within the known backlash wave), and the two new attachments (JevPertus option-scoring on Apertus, a Jev semantic linter) are 2-point same-pattern tail builds from the same HN community — consolidation decay, not periphery expansion into new communities, outlets, or consequential participants. The decisive calibration refutation remains unanswered and neither load-bearing trigger (TypeSafe data-backed rebuttal, frontier typed-decision endpoints) has fired. Heat stays low despite the magnitude-valve and 90th-percentile peer reading: ~11 pts/h vs ~2916 peak and cooling; those are stock measures of the 213-item launch corpus, not current flow.
2026-09-25T13:25:17Z
evidence attached: hn.story.49843708 — Independent adoption of Jev for linting is third-party usage evidence for the significant structured-decisions case.
2026-09-25T13:25:17Z
evidence attached: hn.story.49843906 — Independent Jev-style option scoring on open Apertus shows the structured-decision pattern spreading beyond TypeSafe's own model.
2026-09-25T08:42:38Z
No meaning change. The new attachment is a one-point routing anecdote restating the already-priced pattern — frontier models win on quality, Jev earns its keep only on micro-judgments — and the mild rate uptick (~17 vs ~11 pts/h, flagged 'accelerating') is small-post cadence from the same HN/r/LocalLLaMA communities, not periphery expansion into new communities, outlets, or consequential participants; the magnitude-valve reading remains a stock measure of the 211-item launch corpus, the calibration refutation stays unanswered, and neither load-bearing trigger (TypeSafe data-backed rebuttal, frontier typed-decision endpoints) has fired.
2026-09-25T08:24:29Z
evidence attached: reddit.post.1wpq62p — Independent user assessment of Jev vs Claude quality plus a concrete routing-adoption report bears directly on the 'frontier-comparable' claim.
2026-09-25T06:44:51Z
Meaning unchanged. The velocity_spike resolves to the Flick LinkedIn computer-use post's engagement tail — an artifact already assimilated into the assessment — and the two new attachments (Jevper wire-format client, Knowledge Signal rubric tool) are one-to-two-point tail-end instances of the already-priced bounded-adoption pattern from the same HN community, not periphery expansion into new communities or consequential participants. The decisive calibration refutation remains unanswered by TypeSafe and neither load-bearing trigger (data-backed rebuttal, frontier typed-decision endpoints) has fired. Heat stays low despite the magnitude-valve reading: ~11 pts/h vs ~2898 peak and cooling; the 96.4 percentile is a stock measure of the 210-item launch corpus, not current flow.
2026-09-25T06:25:20Z
evidence attached: hn.story.49840561 — First visible third-party artifact built on TypeSafe's Jev model — early external-adoption evidence contextualising whether Jev becomes a real structured-decision substrate.
2026-09-25T06:25:20Z
evidence attached: hn.story.49840721 — First independent client constrained to the Jev wire format — near-zero engagement but genuine earliest ecosystem evidence for whether Jev becomes a real interface.
2026-09-25T05:26:04Z
No meaning change. The two new attachments are one-point same-pattern builds from the same HN community (a Jev-based PR-review tool, Perch semantic linting) — tail-end instances of the already-priced bounded-adoption pattern, not periphery expansion into new communities, outlets, or consequential participants; the calibration refutation remains unanswered and neither load-bearing trigger (TypeSafe data-backed response, frontier typed-decision endpoints) has fired. Heat stays low despite the magnitude-valve reading: 10.3 pts/h vs ~2890 peak and cooling is consolidation decay — the 96.7 percentile is a stock measure of the 208-item launch corpus, not current flow.
2026-09-25T05:23:37Z
evidence attached: hn.story.49840204 — Third-party artifact applying Jev to semantic code linting is independent adoption evidence for the Jev structured-decisions case, beyond TypeSafe's own claims.
2026-09-25T05:23:36Z
evidence attached: hn.story.49840300 — Third-party open-source code-review tool explicitly built on Jev is early ecosystem-adoption evidence for the significant Jev structured-decisions case.
2026-09-25T03:40:09Z
The 'Calibration Beats Accuracy' essay is pro-Jev philosophy, not measurement: it offers no data against the jevals/fair-die refutation, its thread flags likely astroturf (coordinated aged accounts), and a commenter restates the load-bearing point that calibration must be proven per-dataset — the refutation remains unanswered. Current flow is still 1–2-point same-pattern posts from the same two communities at ~99% below peak, so the magnitude-valve spread reading remains stock, not continuing periphery expansion; meaning unchanged and neither load-bearing trigger (TypeSafe data-backed response, frontier typed-decision endpoints) has fired.
2026-09-25T03:25:49Z
evidence attached: hn.story.49839510 — Independent technical analysis arguing calibration-over-accuracy for Jev/System One models directly contextualises the Jev structured-decisions case's core calibration claim.
2026-09-25T00:10:38Z
No meaning change. The four new attachments each instantiate already-priced patterns from the same two communities — a fourth Jev-as-Claude-Code-router, a Jev-inspired multimodal derivative (Valen), a browser-subtask tool, and one implementer's jev-1.13 deployment note via OpenRouter (notable mainly for confirming TypeSafe API slots remain exhausted) — and the velocity_spike is a +8-point uptick on the already-priced Flick demo. Per the prior look's explicit condition, heat drops to low: the magnitude-valve spread reading is a stock measure of the launch-window corpus (205 items, ~99% below peak, cooling) rather than continuing periphery expansion — no new community, outlet, or consequential participant has joined, and the load-bearing triggers (TypeSafe rebuttal of the calibration refutation, frontier-lab typed-decision endpoints) have not moved.
2026-09-24T23:33:57Z
evidence attached: hn.story.49838009 — Third independent Jev-based tool in one batch (browser sub-tasks at ~1/3 cost), an adoption echo across unrelated projects.
2026-09-24T23:33:57Z
evidence attached: hn.story.49838068 — A Jev-inspired multimodal decision model shows the pattern spawning derivatives, relevant to whether Jev becomes a category.
2026-09-24T23:33:56Z
evidence attached: hn.story.49837537 — Third-party tool using Jev to pick Claude Code's effort/model/skill per prompt is direct adoption evidence for the Jev-as-primitive case.
2026-09-24T23:33:56Z
evidence attached: reddit.post.1wpfu5j — Independent implementer benchmarking jev-1.13 via OpenRouter with automated Opus/Astra/Sol evaluation — early third-party usage evidence the significant Jev case needs.
2026-09-24T22:53:34Z
Meaning unchanged: the only new item — a 1-point 'JEV-Star' StarCraft II control demo — is one more thin same-community game demo in the already-priced saturated open ecosystem, and the engagement upticks (astroturf complaint ~1.16k, 'isn't new tech' 721) re-litigate settled ground without new data. The case holds its post-refutation posture (fast/cheap bounded classifier, uncalibrated confidence); medium heat now rests on Scott's active Venture World stake and pending triggers rather than volume, and should drop to low if the next look is again this thin.
2026-09-24T20:37:40Z
evidence attached: hn.story.49834465 — Distinctive JEV naming plus low-cost structured-control framing suggests possible unconfirmed third-party adoption of the Jev approach worth flagging for the watching case.
2026-09-24T18:12:10Z
No meaning change: the three new items (Cbjev, sub-30-min Mac Air fine-tune guide, Jev-vs-cross-encoder) are thin same-community additions confirming the already-priced saturated-open-field picture, and the RLCD explainer's comment uptick (55 pts) re-litigates a question already settled against TypeSafe without producing new data. The case holds at its post-refutation posture — fast/cheap bounded classifier with uncalibrated confidence — and now waits on TypeSafe's response to the calibration refutation or a frontier-lab typed-decision endpoint for its next meaning change.
2026-09-24T16:38:17Z
evidence attached: hn.story.49830413 — Independent open-source one-pass typed-decision engine is early replication evidence for the structured-decision model class the significant Jev case claims.
2026-09-24T16:38:17Z
evidence attached: hn.story.49831831 — A sub-30-minute Mac Air fine-tune guide is accessibility evidence for the claim that Jev-class structured-decision models are dramatically cheaper than autoregressive LLMs.
2026-09-24T16:38:16Z
evidence attached: hn.story.49831937 — shared external link with case evidence
2026-09-24T15:29:40Z
The calibration claim — RLCD's load-bearing differentiator — moved from 'contested' to 'refuted on ground truth': an independent fair-die test with published code/data (400 hidden rolls) shows Jev returning ~83% on face 1 every time where truth is 1/6 (~19% accuracy), converging with jevals' calibration-gap losses and an independent can't-be-calibrated analysis. Jev's meaning narrows again: a fast, cheap, bounded state-derived classifier whose confidence outputs are uncalibrated scores — unusable for risk-gated automation or escalation thresholds without per-task recalibration — while the OpenJudgement-4B add is one more entry in the saturated open-replacement field and the CLM velocity spike is already-priced backlash.
2026-09-24T15:25:47Z
evidence attached: reddit.post.1wp2xpn — Early open-weight replication attempt of the Jev pattern contextualizes whether the approach is replicable, though it claims no parity (not corroboration).
2026-09-24T15:25:47Z
evidence attached: hn.story.49830385 — shared external link with case evidence
2026-09-24T14:41:07Z
The only new item since the last look — a third-party RLCD explainer at negligible traction (2 pts, 0 comments, content unseen) — adds one more marginal entry to the already-saturated third-party dissection of Jev's technical claims and resolves neither the 'is RLCD actually RL' dispute nor the calibration question. The case's meaning is unchanged: significant but narrow, real-but-rapidly-commoditizing niche in backlash/consolidation; medium heat is held on continued slow periphery accretion and Scott's active Venture World stake, not new substance.
2026-09-24T13:28:09Z
evidence attached: hn.story.49829625 — Third-party technical explainer of the training method behind Jev, material context for re-judging the significant Jev claims case.
2026-09-24T12:46:10Z
The newest substantive item is a third-party in-the-wild computer-use demo — Jev driving real UI automation on a live Chrome profile in ~7s alongside ChatGPT/Astra harnesses — the first non-game confirmation of the exact real-time-automation promise, which strengthens the defensible niche without touching the dented calibration/parity story; CLM's same-day Reddit+HN echo is one more entry in the already-saturated open-replacement field. The case's meaning holds at 'narrow, real, rapidly commoditizing': medium heat is earned by slow but continuing periphery accretion (new implementations, second communities, new use classes) and Scott's active Venture World stake, while the big backlash threads remain already-priced repetition, and the ~97%-decayed velocity argues against anything louder.
2026-09-24T12:24:47Z
evidence attached: reddit.post.1woyzve — Apparent real-world use of Jev driving real-time computer-use automation in seconds — third-party adoption evidence bearing on the Jev low-latency real-time-automation claim.
2026-09-24T07:51:28Z
The new CLM evidence (API-parity projection head on Qwen3-8B, echoed on Reddit and HN) is confirmatory, not new: it adds one more entry to an already-saturated open-replacement field, and its HN echo has no traction — so the case's meaning is unchanged. What did move is the attention phase: the episode's largest live threads are now anti-Jev (1k+-point astroturfing complaint, 576-point 'not new tech' post), pricing TypeSafe's community standing as souring while the substantive open questions (calibration, frontier-parity, native frontier endpoints) stay where the last look left them. Holding medium despite the 99th-percentile spread reading: velocity is ~3% of peak and the fast movers are meta-backlash threads whose content the case has already priced, not new reach.
2026-09-24T07:24:45Z
evidence attached: hn.story.49827145 — Same CLM-vs-Jev story surfacing on HN — second independent community carrying the episode within the window.
2026-09-24T07:24:45Z
evidence attached: reddit.post.1wouby6 — Open-weights CLM claims full API-level parity with Jev with real Reddit traction — a direct challenge to whether Jev's structured-decision niche holds.
2026-09-24T06:39:33Z
Nine days in, the episode has flipped from expansion to consolidation: converging independent results (jevals calibration gaps favoring frontier LLMs, a near-chance objectively-labeled market screener on a Jev-style harness, 25-line replications and a saturated open-source field, plus an 'OpenAI absorbs the pattern' thesis) recast Jev from extraordinary unverified claims to a real but narrow, rapidly commoditizing typed-decision layer whose calibration and parity claims are actively dented. Heat cools to medium — the 99th-percentile cohort speed masks a ~98% decay from peak velocity, new periphery additions have thinned to 0–2-point posts, and the newest big threads are meta-backlash (astroturf complaints, 'not new tech') rather than new reach — while the case stays significant for Scott's active Venture World integration.
2026-09-24T06:24:08Z
evidence attached: reddit.post.1wot9j5 — Independent practitioner builds a Jev-style structured screener on Qwen3.8-27B and finds near-chance selection against market labels despite convincing explanations — rare objectively-labeled context for the structured-judgment calibration claims.
2026-09-24T01:05:04Z
grounded: converges/high — Jev converges with Scott’s Micro-Judgement Pattern and risk-based triage doctrine by supplying cheap typed judgments for routing, validation, and escalation; bo
2026-09-24T01:01:51Z
Independent bounded deployments now show that Jev can materially reduce latency and cost for verifier and bulk-classification workloads, making the decision-model pattern significant for Scott even though TypeSafe’s headline parity and calibration claims remain unproven. Heat stays high despite cooling momentum because implementations and competing runtimes continue expanding across communities at a 98.9th-percentile attention rate.
2026-09-24T00:31:35Z
evidence attached: hn.story.49823839 — Third-party experiment using Jev for judgment-based guardrails is early adoption evidence for whether Jev becomes a real decision-engine substrate.
2026-09-24T00:31:35Z
evidence attached: hn.story.49824306 — Skeptical benchmark pushback on Jev hype is contradicting/contextualising evidence for the accelerating Jev claims case.
2026-09-24T00:31:35Z
evidence attached: reddit.post.1wol78y — Independent third-party (Berman viral demo) corroboration of JEV's structured-decision cost/latency claims at scale, plus pipeline details clarifying how the 724-ad/216ms result was achieved — the most valuable kind of attach.
2026-09-23T21:46:36Z
evidence attached: hn.story.49821631 — Third-party (Tessl) measured use of Jev verifiers at 13.6x speed and 2.7x lower cost vs GPT Luna 6 independently corroborates the structured-decision economics claim on an accelerating case.
2026-09-23T21:46:36Z
evidence attached: hn.story.49821894 — Independent (if informal) third-party benchmark of Jev's judgment quality against LLMs bears directly on the accelerating case's frontier-comparable-decisions claim.
2026-09-23T21:46:36Z
evidence attached: reddit.post.1woameh — Third-party effort applying Jev-style decision models to agent evaluation is independent ecosystem-uptake evidence for the decision-model pattern beyond TypeSafe's own claim.
2026-09-23T21:46:35Z
evidence attached: reddit.post.1woe70t — Substantive skeptic critique that Jev's impressive comparisons use the wrong baseline (LLMs vs classic zero-shot classifiers) — material counter-evidence for weighing TypeSafe's accelerating claims.
2026-09-23T21:46:35Z
evidence attached: reddit.post.1wohvds — Independent round-2 benchmark re-runs Jev against SemIf, Laya, Von and a general 4B model on a fresh reproducible draw — direct outside evidence on the case's performance claim.
2026-09-23T21:46:35Z
evidence attached: reddit.post.1woie1t — Third-party Go adapter implementing the Jev decision-API schema on llama.cpp shows the Jev pattern spreading beyond TypeSafe's own model into the local ecosystem.
2026-09-23T16:27:31Z
evidence attached: hn.story.49817558 — Third-party TUI and decision-feedback tooling around Jev is ecosystem-formation evidence relevant to judging the case's trajectory.
2026-09-23T16:27:31Z
evidence attached: hn.story.49817923 — TypeSafe co-founder's insider account of why Jev could not be built at OpenAI is direct first-party context for this case.
2026-09-23T16:27:31Z
evidence attached: hn.story.49816899 — Independent technical critique directly challenging the calibrated-probabilities claim at the core of this accelerating case.
2026-09-23T15:28:49Z
evidence attached: hn.story.49817395 — Third-party post examining Jev's determinism/probabilistic behavior — independent attention to the accelerating Jev structured-decisions case.
2026-09-23T15:28:49Z
evidence attached: hn.story.49816972 — shared external link with case evidence
2026-09-23T14:27:57Z
evidence attached: hn.story.49815958 — Claim that open models beat Jev on accuracy, speed, and cost is direct adversarial evidence against Jev's value proposition in an accelerating case.
2026-09-23T14:27:57Z
evidence attached: reddit.post.1wo6o0f — High-engagement community allegation of VC-funded Jev astroturfing materially contextualizes how much of Jev's accelerating momentum is paid marketing rather than organic adoption.
2026-09-23T14:27:57Z
evidence attached: reddit.post.1wo6x7e — Third-party local Jev-like system (Bonsai-Llama-Jev) independently benchmarked against Jev-1.13 shows the Jev structured-decision approach spreading into local inference.
2026-09-23T10:22:03Z
evidence attached: reddit.post.1wo1k94 — This released open-source decision engine is a directly relevant alternative to Jev, but its local-model performance claims remain unverified.
2026-09-23T09:22:30Z
evidence attached: reddit.post.1wo0a9t — The local proxy implements the Jev protocol and describes a concrete fallback-and-drift workflow, materially contextualizing the case's low-cost structured-decision path.
2026-09-23T09:22:29Z
evidence attached: reddit.post.1wo07r8 — shared external link with case evidence
2026-09-23T08:21:46Z
evidence attached: hn.story.49812769 — The post offers a compact implementation artifact for Jev, relevant technical context for the open claim about fast structured decisions, but not independent performance validation.
2026-09-23T07:22:41Z
evidence attached: hn.story.49812519 — The released conversion skill offers an early workflow example for migrating structured agent calls to Jev, though it is a creator's unvalidated report rather than independent corroboration.
2026-09-23T02:21:41Z
evidence attached: reddit.post.1wnt7kv — The local stuntd release is a concrete Jev-compatible deployment and learning path that bears on the ecosystem’s cost and latency claims.
2026-09-23T02:21:41Z
evidence attached: reddit.post.1wnt9ae — This describes applying Jev-style typed decisions to response, RAG, and agent-trace evaluation, contextualizing its intended use.
2026-09-23T02:21:41Z
evidence attached: reddit.post.1wntc8m — Independent task-level tests add evidence about Jev’s practical accuracy against local alternatives, though the sample is small and self-selected.
2026-09-22T21:23:03Z
evidence attached: hn.story.49806864 — An open-source Jev architecture artifact materially contextualizes the emerging structured-decision runtime behind the open case.
2026-09-22T21:23:03Z
evidence attached: reddit.post.1wnmav6 — The reported chunk-free classification and confidence results materially contextualize Jev's positioning as a fast structured-decision and retrieval alternative.
2026-09-22T18:24:11Z
evidence attached: hn.story.49804866 — The released offline TinyJev implementation materially supports the case that Jev-style structured decision models are becoming deployable beyond a hosted service.
2026-09-22T18:24:11Z
evidence attached: hn.story.49805823 — A local Mac implementation of a Jev-like model is direct artifact evidence for cheaper on-device structured decisions.
2026-09-22T18:24:10Z
evidence attached: reddit.post.1wngz6r — A user experiment raises a concrete concern about demographic sensitivity in Jev's probabilistic decisions.
2026-09-22T18:24:10Z
evidence attached: reddit.post.1wngtx3 — The JevBench artifact is direct early evidence about evaluation and missing modality support for the open structured-decision model.
2026-09-22T17:28:45Z
evidence attached: hn.story.49800574 — JevBench directly bears on whether Jev’s claimed structured-decision performance is reproducibly measurable, though it is low-signal coverage.
2026-09-22T17:28:45Z
evidence attached: hn.story.49803758 — Code-review calibration provides a concrete downstream use case for evaluating Jev’s structured-decision quality.
2026-09-22T17:28:45Z
evidence attached: hn.story.49803834 — A browser-agent application is relevant evidence for whether Jev-like typed decisions work beyond the originating product claims.
2026-09-22T17:28:45Z
evidence attached: hn.story.49804080 — A separate pure-Go and SIMD/PTX implementation provides corroborating evidence about the practicality of Jev-like typed decisions.
2026-09-22T17:28:44Z
evidence attached: hn.story.49804788 — Production comparison results provide independent evidence relevant to Jev's claimed latency and quality advantages over cross-encoders.
2026-09-22T16:23:55Z
evidence attached: hn.story.49803473 — The released skill offers implementation evidence around structuring Jev requests for reliable agent decisions.
2026-09-22T16:23:55Z
evidence attached: hn.story.49803197 — This is practical evidence that Jev-style decisions can replace recurring LLM turns in coding agents.
2026-09-22T16:23:55Z
evidence attached: reddit.post.1wnci2c — The independent open-weight reimplementation provides technical context and partial external validation for Jev’s structured-decision approach.
2026-09-22T15:22:43Z
evidence attached: hn.story.49802161 — Directly bears on whether OpenAI is becoming a competitive threat to Jev, though the linked commentary is not independent performance evidence.
2026-09-22T15:22:43Z
evidence attached: reddit.post.1wnbuw3 — A concrete small-model derivative applies Jev-style low-latency structured decisions to image classification, modestly extending the case beyond text automation.
2026-09-22T13:23:24Z
evidence attached: hn.story.49800531 — This is another concrete attempt to turn a small local model into a typed, low-latency decision function.
2026-09-22T13:23:24Z
evidence attached: hn.story.49800787 — Blink is an independent executable artifact exploring Jev-like low-latency typed decisions in C and WASM.
2026-09-22T11:23:45Z
evidence attached: hn.story.49799065 — A production account is useful independent contextual evidence about where Jev works and fails, directly informing the open Jev adoption hypothesis.
2026-09-22T11:23:45Z
evidence attached: hn.story.49799118 — This provides additional evidence that Jev is being used as a target for automated harness optimization, materially informing the open Jev capability and adoption hypothesis.
2026-09-22T10:21:58Z
evidence attached: hn.story.49798734 — This comparison directly bears on Jev's claimed cost, latency, and capability advantage over frontier LLMs.
2026-09-22T09:23:22Z
Another Claude Code router listing extends the implementation ecosystem but supplies no measured savings or routing-quality result, so it does not strengthen Jev’s performance claims. High attention remains warranted by the broad HN/Reddit spread and continuing expansion of implementations, not by the repetitive architecture and priority debate.
2026-09-22T09:21:41Z
evidence attached: hn.story.49798407 — shared external link with case evidence
2026-09-22T05:23:42Z
The latest routing anecdote and technical write-up add incremental adoption and explanatory coverage, but no measured result that changes the assessment. Attention remains high because implementations and use cases continue spreading across communities, while Jev’s comparative accuracy and calibration claims remain unsettled.
2026-09-22T05:22:47Z
evidence attached: hn.story.49796843 — Simon Willison's technical write-up is independent corroboration and context for the Jev model architecture and release.
2026-09-22T05:22:47Z
evidence attached: reddit.post.1wn07tn — User experience with Jev model routing provides adoption evidence for its claimed low-cost model-selection role.
2026-09-22T04:29:31Z
New Jev-versus-Laya and domain-specific experiments reinforce that decision-model performance is workload- and training-data-dependent rather than universally frontier-comparable. They add useful evaluation paths but remain too thinly documented to settle calibration or accuracy, while the open implementation and agent-evaluation periphery continues expanding rapidly.
2026-09-22T04:21:40Z
evidence attached: hn.story.49796639 — A direct Jev-versus-Laya comparison is relevant independent evaluation evidence for the open Jev structured-decision case.
2026-09-22T04:21:40Z
evidence attached: reddit.post.1wmyvzp — A concrete user experiment suggests Jev’s value depends on domain-specific data and calibrated decision labels, materially contextualizing its structured-decision claims.
2026-09-22T04:21:40Z
evidence attached: reddit.post.1wmypc5 — shared external link with case evidence
2026-09-22T03:27:42Z
Kev’s reported Apache-2.0 model family, training code and TypeSafe-compatible local API strengthen the shift from Jev-specific launch hype toward a reproducible deployment pattern. The implementation periphery is still expanding rapidly, sustaining high attention, but neither Kev nor the many demos validates Jev’s frontier-equivalence or calibration claims.
2026-09-22T03:21:58Z
evidence attached: reddit.post.1wmxx9w — The released local Kev models provide an open, runnable implementation and deployment path for the same Jev-like structured-decision approach.
2026-09-22T00:22:49Z
Jevopt extends Jev’s implementation periphery into compiler optimization, but its qualified “sometimes” result lacks enough measurements or methodology to establish an advantage over conventional heuristics. The latest DeepSeek comparison repeats the existing picture—latency and selective coverage may be compelling despite weaker calibration—so broad, continuing experimentation sustains high attention without validating frontier equivalence.
2026-09-22T00:22:32Z
evidence attached: hn.story.49795171 — Jevopt is a concrete compiler-optimization application that tests whether Jev's structured decisions can replace conventional heuristic choices.
2026-09-22T00:22:32Z
evidence attached: reddit.post.1wmt5t9 — This comparison provides adoption and calibration context for Jev's claimed low-latency structured-decision positioning, albeit with minimal engagement.
2026-09-21T23:23:47Z
A newly quoted human-label benchmark challenges Jev’s calibration advantage while reporting stronger decision coverage at a fixed accuracy threshold: useful automation may depend more on selective reliability than universally better probabilities. The underlying methodology remains unavailable here; a moderation integration and technical-interview appearance extend the implementation and discussion periphery, sustaining high attention without establishing frontier parity.
2026-09-21T23:22:20Z
evidence attached: hn.story.49794590 — A CEO interview provides relevant first-party context on Jev's intended production role, but adds no independent validation.
2026-09-21T23:22:20Z
evidence attached: reddit.post.1wms3sr — A small platform's moderation deployment provides early adoption evidence for Jev's structured-decision workflow, though not reliability validation.
2026-09-21T23:22:20Z
evidence attached: reddit.post.1wmre0b — The benchmark report materially challenges Jev's calibration advantage while supporting its claimed high-coverage decision behavior.
2026-09-21T22:42:42Z
The latest Gemma 4 E2B classifier experiment adds another independent builder to the local-alternative ecosystem, but its truncated account provides no measured accuracy, calibration, latency or cost comparison. It does not change the performance assessment; broad top-decile HN/Reddit spread and continuing implementation expansion still justify high attention.
2026-09-21T22:23:12Z
evidence attached: reddit.post.1wmpwtl — A hands-on attempt to reproduce Jev-like fast probabilistic classification with a small local model provides useful independent context on the approach's calibration and cost tradeoffs.
2026-09-21T21:54:15Z
The claimed survey of 600+ Jev builds reinforces the already-established emphasis on classification, routing and scoring, but supplies neither an auditable inventory nor new implementation results in the available excerpt. It does not strengthen the performance hypothesis; high attention remains justified by the broad cross-platform spread and recently expanding implementation ecosystem, not this roundup alone.
2026-09-21T21:22:54Z
evidence attached: reddit.post.1wmp9xw — The 600-build survey supplies independent ecosystem evidence that Jev's low-cost structured decisions are being used for high-volume classification, routing, and scoring.
2026-09-21T20:59:27Z
The latest one-token LLM argument repeats an already-known baseline challenge, without measurements establishing parity with Jev’s parallel decisions or claimed calibration. It changes neither the performance assessment nor the practical testing opportunity; broad cross-platform spread and the recently expanding implementation ecosystem still warrant high attention.
2026-09-21T20:23:24Z
evidence attached: reddit.post.1wmnyuz — The discussion directly challenges the claimed distinctiveness of Jev's one-token structured-decision approach and is relevant to re-judging that case.
2026-09-21T19:37:10Z
A user-run Jev–Laya comparison claims a broad Jev accuracy advantage, but the supplied excerpt stops before the results and exposes neither methodology nor calibration measurements; it does not yet change the performance assessment. The expanding implementation ecosystem and substantial cross-platform spread still warrant high attention, independently of whether Jev's strongest claims hold.
2026-09-21T19:22:48Z
evidence attached: reddit.post.1wmmi5d — A user-run comparison provides independent early evidence about Jev's accuracy and calibration claims against another compact decision model.
2026-09-21T18:23:56Z
A disclosed small retrieval/memory pilot adds a measured ranking result for a separate decision judge, making this more relevant to evaluating judge–writer separation rather than merely replacing classifiers. Its reported AUROC advantage does not establish calibrated probabilities or end-to-end retrieval gains, and the available excerpt does not isolate Jev’s contribution from the second judge.
2026-09-21T18:21:58Z
evidence attached: reddit.post.1wmkr01 — A disclosed small pilot materially tests Jev-style structured decisions, finding better ranking but poor verbal confidence calibration and worse end-to-end retrieval.
2026-09-21T16:32:00Z
New reports extend the application periphery into agent linting and bug-bounty workflows, but their headline-only evidence does not establish implementation quality or the claimed 31% coding-agent speedup. Continued cross-platform spread and expanding applications sustain high attention without strengthening the frontier-equivalence or calibration claims.
2026-09-21T16:23:35Z
evidence attached: hn.story.49788625 — A bug-bounty workflow using TypeSafe AI provides practical adoption context for the open Jev case.
2026-09-21T16:23:35Z
evidence attached: hn.story.49788876 — This is direct adoption evidence for Jev as a low-cost structured-decision component in an agent harness.
2026-09-21T16:23:35Z
evidence attached: reddit.post.1wmfwk0 — The linked hands-on report provides early deployment evidence about Jev's latency and coding-agent speed benefits.
2026-09-21T15:48:02Z
New documented local implementations make the alternative-backend comparison more actionable: Laya MPS reports explicit memory/latency trade-offs, while OpenDecision offers local typed decisions and document evidence interfaces. This sustains high attention as the implementation ecosystem expands, but neither the headline-only tool-call test nor the constrained text-generation experiment establishes Jev’s calibration or frontier equivalence.
2026-09-21T15:41:04Z
evidence attached: hn.story.49787404, hn.story.49787265, hn.story.49787050 — The two local Jev-style implementations and a Jev-controlled game demonstration are independent evidence that TypeSafe’s typed-decision interface is already prompting compatible tools and applications.
2026-09-21T15:25:01Z
evidence attached: hn.story.49788402 — Independent testing of Jev on 100 real agent tool calls materially bears on whether its typed decisions work in practical agent automation.
2026-09-21T15:25:00Z
evidence attached: hn.story.49787047 — A 521-model Jev ensemble provides independent ecosystem evidence relevant to whether Jev-style typed decisions are becoming a practical alternative to ordinary LLM generation.
2026-09-21T14:20:05Z
The implementation periphery continues expanding across developer tools, model routing, grading, local alternatives and a Vercel Labs template, sustaining high attention without validating Jev’s calibration or frontier-equivalence claims. The latest CLI and discussion mostly reinforce established availability and ecosystem momentum rather than changing the technical assessment.
2026-09-21T13:22:10Z
evidence attached: hn.story.49786725 — A released CLI wrapper provides a concrete client integration for the existing Jev typed-decision runtime.
2026-09-21T10:22:29Z
evidence attached: reddit.post.1wm8btm — Directly questions the central calibration claim behind Jev and supplies useful independent scrutiny despite limited engagement.
2026-09-21T10:22:29Z
evidence attached: reddit.post.1wm83ts — It usefully explains the boundary between Jev's cheap typed decisions and Claude's remaining free-form reasoning work.
2026-09-21T09:22:28Z
evidence attached: hn.story.49784690 — A Vercel Labs template is independent adoption evidence that Jev is being integrated into application workflows.
2026-09-21T09:22:28Z
evidence attached: hn.story.49784831 — A concrete Jev grading tool provides ecosystem evidence for typed, rule-governed model decisions.
2026-09-21T08:21:38Z
evidence attached: hn.story.49784390 — A concrete Jev-based trading-engine artifact provides early adoption evidence for Jev-style typed decision systems.
2026-09-21T07:23:02Z
evidence attached: hn.story.49783694 — The JEV agent-monitoring discussion materially contextualizes the open case about JEV as a low-latency control and decision system.
2026-09-21T07:23:02Z
evidence attached: hn.story.49783999 — The open Kev repository provides an independent, inspectable Qwen-based implementation of the same tiny typed-decision-model direction, making the existing Jev hypothesis more testable.
2026-09-21T06:21:24Z
evidence attached: hn.story.49783502 — The released audit tool is a concrete artifact testing where Jev-like low-latency structured decisions can replace ordinary code logic.
2026-09-21T04:21:46Z
evidence attached: reddit.post.1wm2fpf — JevGraph provides an applied knowledge-graph extraction result that materially tests Jev's claimed speed and cost advantages against general LLMs.
2026-09-21T03:22:09Z
evidence attached: hn.story.49781694 — The article provides relevant technical context on Jev’s probabilistic-logic positioning and relationship to agent decision systems.
2026-09-21T03:22:09Z
evidence attached: reddit.post.1wm0cn0 — A released open-weight, non-autoregressive typed-decision model is directly relevant comparison and corroboration for the structured-decision alternative to autoregressive LLMs.
2026-09-21T02:21:38Z
evidence attached: hn.story.49782061 — This is direct ecosystem context for Jev and supports the open case on its System-1 agent architecture.
2026-09-21T01:21:45Z
evidence attached: hn.story.49781612 — A released local alternative directly tests whether Jev-like typed, low-latency decisions can be delivered without the hosted system.
2026-09-20T23:21:41Z
evidence attached: hn.story.49780849 — This released Jev evaluation project directly supports the open hypothesis that typed Jev decisions can replace LLM judges in low-latency evaluation workflows.
2026-09-20T22:21:55Z
evidence attached: reddit.post.1wlu9rd — A small-scale DIY reproduction supports the broader hypothesis that typed or probability-based decisions can be built from ordinary open models without a dedicated Jev model.
2026-09-20T18:22:51Z
evidence attached: hn.story.49778191 — This independent CLINC150 evaluation directly tests whether Jev delivers reliable low-latency structured decisions in a realistic intent benchmark.
2026-09-20T18:22:50Z
evidence attached: hn.story.49778273 — The released Atari integration provides additional evidence about Jev's practical scope beyond closed-set classification.
2026-09-20T18:22:50Z
evidence attached: hn.story.49778162 — A public Jev chatbot artifact provides independent evidence about how the structured-decision model behaves when adapted for free-form dialogue.
2026-09-20T18:22:50Z
evidence attached: reddit.post.1wloev1 — Hands-on testing supports Jev's potential value as a calibrated low-latency routing and mode-selection component rather than merely a classifier.
2026-09-20T17:24:50Z
New local-inference and fine-tuning reports extend the implementation ecosystem, but do not establish cheaper reproduction of Jev’s capabilities: the Qwen developer explicitly says the model is nowhere near Jev, and used DeepSeek—not Jev—for synthetic training data. The supplied laya.cpp excerpt does not substantiate the attachment’s ggml/CUDA and API claims; the separate Jev-as-teacher report remains title-only.
2026-09-20T17:23:40Z
evidence attached: hn.story.49777639 — Using Jev to teach a smaller model provides adoption context for its claimed role in low-cost structured decision workflows.
2026-09-20T17:23:40Z
evidence attached: reddit.post.1wllv4i — An independently released Qwen3.5-4B fine-tune provides useful evidence that Jev-like typed decisions can be reproduced and compressed into smaller open models.
2026-09-20T17:23:40Z
evidence attached: reddit.post.1wlmkm9 — A usable ggml/CUDA implementation and API endpoint materially strengthen the case that Jev-style structured decisions can reach practical local inference.
2026-09-20T16:23:03Z
New reports extend experimentation into canned-response conversation, agent-rule gates and simulated robot fleets, sustaining high attention without materially strengthening the performance hypothesis. Attachment summaries overstate the evidence: the M4/CoreML claim concerns Laya, not Jev, and the compiler and robot-cost reports remain title-only, with no established first-party provenance or inspected results.
2026-09-20T16:22:20Z
evidence attached: hn.story.49776827 — A concrete simulated-robot deployment and per-million decision cost materially contextualise Jev's claimed low-cost structured decision capability.
2026-09-20T16:22:20Z
evidence attached: hn.story.49777057 — The released Jev compiler is a relevant first-party artifact showing executable gates for rules that agents repeatedly violate.
2026-09-20T16:22:20Z
evidence attached: hn.story.49777106 — Concrete local deployment evidence reports Jev running on an M4 through CoreML at 45 decisions per second.
2026-09-20T16:22:20Z
evidence attached: reddit.post.1wljyk0 — This concrete implementation shows Jev being used for probabilistic intent and policy decisions without token generation, providing practical workflow context for the open case.
2026-09-20T14:27:43Z
A title-only report of a Jev-based OpenRouter router for pi extends the implementation footprint but does not substantiate its claimed Pareto optimality or improve evidence for reliable routing. Continued implementation expansion across broad HN/Reddit activity sustains high attention, while calibration and frontier-equivalence claims remain unresolved.
2026-09-20T14:22:09Z
evidence attached: hn.story.49775968 — This is a concrete router artifact built around Jev, providing direct ecosystem evidence for its model-routing and structured-decision claims.
2026-09-20T13:40:26Z
New title-only reports of agent-testing checks and automatic PR approval extend Jev’s implementation periphery into coding-workflow gates, but supply no evidence of reliable defect detection or safe approval. Continued expansion across substantial HN/Reddit activity sustains high attention; it does not strengthen the calibration or frontier-equivalence claims.
2026-09-20T13:22:11Z
evidence attached: hn.story.49775144 — Automated pull-request approval is a concrete workflow application of Jev's low-latency structured decisions.
2026-09-20T13:22:11Z
evidence attached: hn.story.49775566 — A concrete Jev-powered testing tool provides adoption evidence for the open case's structured-decision model.
2026-09-20T12:26:48Z
The latest provenance question and explainer thread recycle unresolved claims rather than supply architectural evidence; Dethrone adds a reported public-play evaluation surface, but no inspected results. Broad HN/Reddit spread and continuing implementation expansion sustain high attention without strengthening the frontier-equivalence or calibration claims.
2026-09-20T12:21:56Z
evidence attached: reddit.post.1wlf0s2 — A public game provides a concrete demonstration and informal evaluation of Jev's decision-only interface.
2026-09-20T12:21:56Z
evidence attached: reddit.post.1wleg4w — The discussion provides adoption context and clarifies Jev's decision-oriented interface rather than a conventional text-generation model.
2026-09-20T12:21:56Z
evidence attached: reddit.post.1wlfmgq — The attribution concern materially contextualizes Jev's provenance and could affect confidence in Typesafe's claimed decision-model novelty.
2026-09-20T11:22:03Z
A technically described local Qwen decision-head prototype and newly announced SQL integrations extend the implementation periphery beyond agent and game demos, sustaining high attention without validating Jev’s comparative performance. The latest benchmark criticism largely repeats the already-disclosed use of model-generated reference probabilities; it is not a new empirical disproof.
2026-09-20T11:21:18Z
evidence attached: hn.story.49774406 — The DuckDB extension is a concrete integration artifact that makes Jev's typed decisions usable in data workflows.
2026-09-20T11:21:18Z
evidence attached: hn.story.49774592 — A concrete MySQL plugin built on Jev demonstrates an early application path for structured decision models and is especially relevant to the deliberately hunted Jev query.
2026-09-20T11:21:18Z
evidence attached: reddit.post.1wldgdt — The post materially challenges Jev's benchmark validity by identifying model-generated reference labels and agreement-based scoring rather than independently grounded accuracy.
2026-09-20T11:21:18Z
evidence attached: reddit.post.1wle746 — A small open-weight implementation provides practical evidence that Jev-style non-autoregressive decision heads can be prototyped locally, while its limited training prevents validating broad capability.
2026-09-20T10:26:41Z
A new AI Gateway report adds a concrete paid-adoption claim, moving the traction evidence beyond demos, although the supplied HN excerpt does not authenticate Vercel attribution or establish sustained production use. The classical-ML comparison is only a title-level lead; expanding distribution and implementations sustain high attention without validating frontier equivalence or calibration.
2026-09-20T10:22:18Z
evidence attached: hn.story.49774164 — Vercel’s first-party adoption claim materially supports the open case that Jev is gaining real deployment traction.
2026-09-20T10:22:18Z
evidence attached: hn.story.49774364 — This independent comparison provides early contextual evidence about Jev's task-dependent performance rather than universal superiority.
2026-09-20T09:25:43Z
The Street Fighter 2 report adds another structured-state game integration, but supplies no measured result that changes the assessment of Jev’s performance or economics; its claim of no game-specific training is unverified. Broad cross-platform spread and a still-expanding implementation ecosystem sustain high attention, while the latest comments add neither substantive corroboration nor credible contradiction.
2026-09-20T09:21:27Z
evidence attached: reddit.post.1wlbji2 — A concrete real-time game-control deployment provides practical evidence about Jev's low-latency structured decision capability.
2026-09-20T08:25:57Z
A developer reports releasing installable Jev skills across Codex, Claude Code and OpenCode for checkpoints, recovery/escalation and triage, extending integration work from demonstrations toward reusable agent-loop recipes. This is a modest tooling advance, not measured reliability improvement; the expanding implementation ecosystem and broad cross-platform spread sustain high attention while the performance hypothesis remains unsettled.
2026-09-20T08:21:41Z
evidence attached: reddit.post.1wlb1jj — The released cross-harness skills provide concrete integration and reliability recipes that materially support Jev adoption in agent loops.
2026-09-20T00:23:04Z
A developer now reports running Jev alongside three local alternatives in Doom, adding a concrete comparative implementation rather than another substitute announcement. The supplied evidence contains no comparative results, and preselected episodes cannot establish superiority or calibration; the still-expanding implementation ecosystem sustains high attention.
2026-09-20T00:22:25Z
evidence attached: reddit.post.1wl1yzq — The Doom experiment provides concrete local-agent evaluation evidence for Jev's structured action decisions, though it is only a small informal test.
2026-09-19T23:23:03Z
Jev-align adds a title-only lead for adapting decisions to user judgment, not evidence that Jev’s probabilities are calibrated or that adaptation improves held-out accuracy. The expanding tooling periphery and broad cross-platform spread sustain high attention, while the core performance claim remains unverified.
2026-09-19T23:21:34Z
evidence attached: hn.story.49770872 — The released calibration CLI provides practical evidence about making Jev's structured decisions better match user judgment.
2026-09-19T21:46:13Z
Odyssey adds a developer-reported bounded-action simulation, but its trajectory and maneuver timing come from the harness; it is not evidence of autonomous flight planning or calibrated decision quality. Scott’s up-vote reinforces the case’s practical interest, while broad cross-platform spread and continuing implementation activity sustain high attention without strengthening the frontier-equivalence claim.
2026-09-19T21:22:33Z
evidence attached: hn.story.49769916 — A live browser simulation provides independent usage evidence that Jev can make bounded decisions with exposed probabilities, directly contextualising the open structured-decision claim.
2026-09-19T20:23:41Z
The drug-discovery validation-gate report extends the evaluation periphery into another agent workflow, but the supplied title contains no implementation details or results and cannot establish effective validation. Broad cross-platform spread and continuing integration activity still warrant high attention; the evidence for calibrated, matched-quality savings has not improved.
2026-09-19T20:22:14Z
evidence attached: hn.story.49769496 — The report provides an early external test of Jev as a fast validation gate for a concrete agent workflow.
2026-09-19T19:28:56Z
The JSON-Schema adapter announcement adds an integration lead, but its title alone does not establish working compatibility or materially improve the performance evidence. Broad HN/Reddit spread and a still-expanding implementation ecosystem continue to warrant high attention, while matched-quality economics and calibration remain unproved.
2026-09-19T19:21:48Z
evidence attached: hn.story.49769277 — A concrete JSON-Schema adapter materially extends Jev's structured-decision workflow and is relevant to evaluating its integration surface.
2026-09-19T18:31:20Z
S1Code adds a title-only coding-agent implementation lead, extending an already broad ecosystem without establishing better automation outcomes. New criticism and a reported confidently wrong Laya answer sharpen the need to test alternatives independently, but neither refute Jev nor validate its claims; continued cross-community implementation spread still warrants high attention.
2026-09-19T18:21:41Z
evidence attached: hn.story.49768651 — S1Code provides an independent adoption signal for Jev in a decision-first coding-agent harness.
2026-09-19T18:21:41Z
evidence attached: reddit.post.1wks5ip — This user report directly bears on whether Jev is a substantive low-latency alternative to autoregressive LLM decisions, though the enthusiasm is weak evidence.
2026-09-19T17:24:58Z
A code-linked PR-review comparison reports a 1.93× speedup over Luna, reinforcing the existing picture of modest practical latency gains rather than validating launch-scale advantages; its isolated cost figure establishes neither relative savings nor matched quality. This does not materially change the assessment, while the broad, still-expanding implementation ecosystem continues to warrant high attention.
2026-09-19T17:23:18Z
evidence attached: hn.story.49767987 — Directly provides additional discussion and use-case evidence for Jev's structured-decision approach.
2026-09-19T17:23:17Z
evidence attached: reddit.post.1wkr5lz — The reported Jev speed and cost comparison provides early workload evidence for the open case's structured-decision efficiency claims.
2026-09-19T16:22:48Z
The implementation ecosystem is extending beyond hosted text decisions: Von supplies concrete CPU-model release links, while VisionLaya describes an image-input adaptation and Cua introduces a computer-use variant. This materially broadens Scott’s comparison set and sustains high attention, but neither these announcements nor title-only claims of Vercel/Cloudflare uptake validate Jev’s calibration or frontier-equivalent economics.
2026-09-19T16:22:06Z
evidence attached: hn.story.49767430 — This released vision variant independently extends Jev-style calibrated typed decisions from text to image inputs.
2026-09-19T16:22:06Z
evidence attached: hn.story.49767564 — This released CUA-S1 artifact applies the System One typed-decision approach directly to computer-use actions.
2026-09-19T16:22:06Z
evidence attached: reddit.post.1wkoym5 — Independent coverage of Jev's claimed 100x decision-cost reduction and early Vercel/Cloudflare uptake directly bears on the open case.
2026-09-19T16:22:06Z
evidence attached: reddit.post.1wkpxn6 — A released 395M CPU model positioned as a drop-in JEV competitor materially broadens evidence about the low-cost structured-decision model landscape.
2026-09-19T13:30:38Z
New Qwen2.5-0.5B and DiffusionGemma submissions extend the local-alternative periphery, but title-only evidence does not establish working replication or improve the performance case. Broad HN/Reddit spread and continuing implementation activity still warrant high attention; the use-case collection and latest benchmark comments add no inspected results.
2026-09-19T13:21:56Z
evidence attached: hn.story.49766255 — A local Jev implementation backed by DiffusionGemma materially supports the portability and local-deployment dimension of the open structured-decision model episode.
2026-09-19T13:21:56Z
evidence attached: hn.story.49766295 — A Qwen2.5-0.5B Jev-like implementation is useful ecosystem evidence that structured-decision models are being reproduced beyond the original project.
2026-09-19T13:21:56Z
evidence attached: reddit.post.1wklfse — The linked collection of hundreds of Jev use cases provides adoption context for the existing structured-decision model episode, though not independent performance validation.
2026-09-19T12:23:22Z
A new user reports Jev making existing non-reasoning Luna/Gemini workloads 10× cheaper and 2× faster, modestly strengthening the practical efficiency case beyond comparisons with reasoning-heavy baselines, though accuracy and accounting remain unreported. Laya extends the open-alternative discussion but comes from an already represented developer, not an independent architectural replication; broad cross-platform activity still warrants high attention.
2026-09-19T12:21:38Z
evidence attached: hn.story.49765348 — An open-source Jev alternative is independent ecosystem evidence that the structured-decision, low-latency approach is attracting implementation interest.
2026-09-19T09:23:08Z
The new benchmark report introduces a potentially useful labelled-data evaluation lead, but the supplied excerpt contains its setup rather than results; the attachment’s claim of demonstrated latency, cost and calibration advantages is not supported by the visible evidence. Broad cross-platform activity and continuing implementations still warrant high attention, without upgrading confidence in frontier equivalence.
2026-09-19T09:21:37Z
evidence attached: reddit.post.1wkgkrz — The independent benchmark directly supports Jev's hypothesis by reporting lower latency, cost, and better calibration than an autoregressive baseline.
2026-09-19T06:22:16Z
The expanding implementation periphery and substantial HN/Reddit spread now warrant high attention and accelerating ecosystem status, without establishing Jev’s performance claims. The latest Vampire Survivors integration adds another developer-reported experiment, not a measured result; earlier low heat underpriced the breadth of activity.
2026-09-19T06:21:47Z
evidence attached: reddit.post.1wke35a — A concrete public repository uses Jev as a game decision engine, providing practical usage evidence beyond the model's structured-decision claims.
2026-09-19T05:27:32Z
The llama.cpp classifier report adds a concrete local comparison candidate, but repeats the already-established point that closed-set logit scoring is possible with small models. It supplies no matched results or evidence about Jev’s architecture, leaving the quality, calibration and economic claims unchanged.
2026-09-19T05:21:32Z
evidence attached: reddit.post.1wkd1dz — This prior llama.cpp implementation materially contextualizes Jev by showing that logit-based closed-set classification is an established local-model technique rather than a wholly new capability.
2026-09-19T00:30:15Z
The new RLCD discussion raises a question about the undisclosed training method, not evidence that the method is invalid; the accompanying article is title-only. Game-demo comments repeat already-known harness and action-space limitations, leaving Jev a plausible specialized backend without new validation of calibration, frontier-equivalent quality or end-to-end savings.
2026-09-19T00:22:10Z
evidence attached: hn.story.49761730 — Independent article provides additional context for Jev's non-autoregressive structured-decision approach and its positioning.
2026-09-19T00:22:10Z
evidence attached: reddit.post.1wk6iei — Raises a substantive challenge to Jev's RL positioning and clarifies the technical question behind its training method.
2026-09-18T21:07:15Z
A local-model user now reports replacing poorly implemented keyword routing with Jev, adding a relevant deployment anecdote but not a measured advantage over competent routing alternatives. The drone-swarm submission is title-only; neither addition establishes calibration, frontier-equivalent quality or better end-to-end economics.
2026-09-18T20:22:41Z
evidence attached: hn.story.49759706 — A real-time 15-drone simulation is practical usage evidence for Jev's claimed low-latency structured-decision capability.
2026-09-18T20:22:41Z
evidence attached: reddit.post.1wk0st3 — An independent local deployment uses Jev for model routing and post-answer validation, providing concrete workflow evidence for cheap structured decisions.
2026-09-18T18:43:00Z
The new Ruby, semantic-review CLI, yes/no app and Devin-demo submissions broaden the integration leads, but title-only evidence does not establish working implementations or improved economics. The accompanying coverage repeats the launch positioning; none closes the quality, calibration or matched-baseline gap.
2026-09-18T18:22:40Z
evidence attached: hn.story.49757734 — Embedding Jev as a Ruby primitive materially broadens evidence of integration beyond a standalone model API.
2026-09-18T18:22:40Z
evidence attached: hn.story.49757757 — A Jev-powered semantic code-review CLI is a concrete developer-tool use case supporting the case's claim about structured decisions.
2026-09-18T18:22:40Z
evidence attached: hn.story.49757995 — A concrete Devin-built demo provides ecosystem and workflow evidence for Jev's claimed low-latency structured-decision primitive.
2026-09-18T18:22:40Z
evidence attached: hn.story.49758022 — A concrete Jev application provides supporting evidence that its structured yes-or-no decision interface is usable beyond the launch claim.
2026-09-18T18:22:40Z
evidence attached: reddit.post.1wjwjxs — The post independently highlights JEV's structured-decision architecture and real-time automation positioning.
2026-09-18T17:55:48Z
The Pokémon developer now clarifies that a separate LLM loop continually repairs Jev’s harness and that Jev still gets stuck, making the reported progress evidence of a hybrid system rather than standalone decision-model competence. The nine-game cost claim and title-only NanoJev submission add experimentation leads, not validated comparative performance or grounds for renewed urgency.
2026-09-18T17:23:48Z
evidence attached: hn.story.49757421 — The NanoJev repository is a related usable artifact that may extend or operationalize the Jev structured-decision episode.
2026-09-18T17:23:48Z
evidence attached: reddit.post.1wjv6hx — A public interactive Jev demonstration provides practical usage evidence for the structured-decision model beyond its vendor claim.
2026-09-18T17:23:48Z
evidence attached: reddit.post.1wjvh14 — shared external link with case evidence
2026-09-18T16:42:19Z
The Jev-versus-Claude post adds a concrete repository link worth inspecting, but the supplied evidence contains no benchmark results or methodology. It is a validation lead, not yet independent confirmation or contradiction of Jev’s performance claims, so the assessment and urgency remain unchanged.
2026-09-18T16:22:51Z
evidence attached: reddit.post.1wjt8eu — A public Jev-versus-Claude benchmark directly provides independent evidence relevant to Jev's claimed quality and cost advantages.
2026-09-18T14:29:44Z
The Claude Code skill adds another integration claim from the Beads-tool developer, not independent evidence of workflow savings; the supplied excerpt does not show evaluation logs or an inspectable release. The Pong comparison is title-only, so neither addition establishes a performance advantage or materially strengthens the case.
2026-09-18T14:22:31Z
evidence attached: hn.story.49754516 — The Pong demonstration is an independent, though narrow, artifact bearing on Jev's claimed advantage over autoregressive models.
2026-09-18T14:22:31Z
evidence attached: reddit.post.1wjr42f — A released Claude Code skill demonstrates an early practical workflow for routing classification and gating tasks to Jev with logged evaluations.
2026-09-18T12:37:00Z
Ulka adds another experimental browser-agent lead, but the supplied evidence is only an HN title, not an inspected artifact or execution result. It does not materially strengthen the existing implementation evidence or resolve whether Jev improves accuracy-adjusted automation economics.
2026-09-18T12:22:38Z
evidence attached: hn.story.49752776 — A browser-agent artifact provides practical usage evidence for Jev's structured-decision model, albeit without independent performance validation.
2026-09-18T11:33:12Z
A developer now reports building a Jev tool for bulk-labelling Beads issues in a persistent Claude Code workflow, making the application to coding-agent infrastructure more concrete. The supplied excerpt provides no inspectable tool, accuracy results or measured savings, so this remains an implementation lead rather than evidence that Jev improves accepted-decision economics.
2026-09-18T11:22:18Z
evidence attached: reddit.post.1wjmo1i — This is practical usage evidence for Jev replacing autoregressive model calls on a structured coding-workflow task.
2026-09-18T10:28:22Z
OpenJev adds a claimed free hosted deployment of the already-known DiffusionGemma approach, but the supplied excerpt contains neither an accessible deployment link nor execution results; the HN title does not independently verify usability. This is a more specific comparison lead, not corroboration of Jev’s architecture, calibration or frontier-level economics.
2026-09-18T10:21:26Z
evidence attached: hn.story.49752041 — The public OpenJev release independently corroborates that Jev-style typed decision inference is usable outside the originating service.
2026-09-18T10:21:26Z
evidence attached: reddit.post.1wjlyzr — OpenJev is an independent usable implementation of Jev-like typed probabilistic decisions, materially strengthening the case's artifact and deployment evidence.
2026-09-18T09:26:15Z
Probably adds a title-only claim of a Jev-powered workflow language, not an inspected implementation or evidence of improved automation economics. Ecosystem experimentation continues, but this attachment does not strengthen the case for frontier-comparable accuracy, calibration or practical savings.
2026-09-18T09:21:27Z
evidence attached: hn.story.49751902 — A programming language built around Jev is an adoption and artifact signal for the model’s structured-decision workflow.
2026-09-18T07:26:20Z
The earlier prior-art claimant now reports a general-purpose open model and benchmark superiority, but the supplied excerpt exposes neither artifacts nor results, so it remains a comparison lead rather than evidence of Jev reproduction or displacement. The Subway Surfers clip adds another game-control claim without establishing its input pipeline, measured latency or advantage over conventional controllers.
2026-09-18T07:22:19Z
evidence attached: reddit.post.1wjiphh — A public demonstration of Jev performing real-time game control materially contextualizes the model's claimed low-latency decision capability.
2026-09-18T07:22:19Z
evidence attached: reddit.post.1wjieap — This released open model is an independent artifact extending Jev-style non-autoregressive structured decisions and provides useful ecosystem evidence for the open case.
2026-09-18T06:28:32Z
The latest alternatives do not establish local reproduction: one describes a source-available symbolic knowledge-base engine, while the GPU offering supplies only a title. These attachments broaden the comparison set without strengthening evidence for Jev’s calibration, frontier equivalence or practical savings.
2026-09-18T06:21:19Z
evidence attached: hn.story.49750584 — A GPU-runnable local Jev alternative materially supports adoption and portability of the structured-decision approach.
2026-09-18T06:21:19Z
evidence attached: hn.story.49750649 — An open-source TypeSafe alternative is independent ecosystem evidence that Jev-like structured decision inference is becoming reproducible locally.
2026-09-18T03:25:47Z
The new BERT/classifier discussion adds an uninspected typed-decisions benchmark lead, not evidence of Jev’s architecture or comparative performance. It reinforces the need for encoder and reranker baselines without changing the assessment of Jev’s claimed calibration or economics.
2026-09-18T03:21:54Z
evidence attached: reddit.post.1wje4xh — The discussion usefully contextualizes Jev as a typed-decision classifier and raises a concrete comparison with encoder-style models.
2026-09-18T01:38:46Z
Mini-Jev adds another title-only local implementation lead, not an inspected reproduction or evidence that Jev’s capabilities transfer to local LLMs. The expanding experiment list still does not establish calibrated probabilities, frontier-comparable accuracy or matched-workload savings.
2026-09-18T01:21:38Z
evidence attached: hn.story.49748643 — The project is a concrete local implementation of Jev's structured-decision approach and materially informs its portability and practical adoption.
2026-09-18T00:26:44Z
The new odd-number checker and “Talk to JEV” submissions are title-only leads, not inspected artifacts or implementation results; neither validates calibration, architecture or practical savings. They add launch-adjacent experimentation without changing Jev’s standing as a plausible specialized backend awaiting matched-workload validation.
2026-09-18T00:22:50Z
evidence attached: hn.story.49748221 — The linked public JEV artifact materially supports the open case that TypeSafe is offering a non-autoregressive structured-decision system.
2026-09-18T00:22:50Z
evidence attached: hn.story.49747934 — The released Jev example provides concrete artifact evidence for the existing hypothesis about calibrated structured decisions replacing some autoregressive calls.
2026-09-17T22:56:53Z
The reproduction tracker, Doom post and coding-task router are title-only leads, not inspected reproductions or implementation results; Doom was already part of the launch evidence. They broaden the list of possible tests without strengthening the case for calibrated, frontier-comparable decisions or demonstrated routing savings.
2026-09-17T22:22:39Z
evidence attached: hn.story.49747155 — This demo provides a concrete coding-task routing use case for Jev, though it offers little independent evidence beyond the existing artifact.
2026-09-17T22:22:39Z
evidence attached: reddit.post.1wj7v5c — A Doom-control demonstration adds practical workflow evidence for Jev's claimed low-latency structured decisions, though it is not independent validation.
2026-09-17T22:22:39Z
evidence attached: reddit.post.1wj74f4 — A reproduction tracker is directly relevant follow-up evidence for whether Jev's structured-decision claims hold independently.
2026-09-17T21:41:28Z
The Claude Code model-routing submission points directly at Scott’s potential use case, but its title alone establishes neither a usable integration nor routing quality or savings. It adds an implementation lead rather than evidence that changes Jev’s assessment.
2026-09-17T21:22:00Z
evidence attached: hn.story.49746321 — A usable integration routes Claude Code decisions through Jev, providing adoption evidence for the structured-decision model.
2026-09-17T20:46:08Z
The web-form tester retracts a claimed 5000x saving in the post title, but the supplied excerpt contains neither corrected results nor the cause of the benchmark error. This adds a cautionary evaluation lead, not an established finding about form complexity or harness overhead, and leaves Jev’s practical economics unresolved.
2026-09-17T20:23:00Z
evidence attached: reddit.post.1wj3lsw — Independent testing materially qualifies Jev's headline cost claim by exposing form complexity and end-to-end harness overhead while still testing its practical decision-model use.
2026-09-17T19:23:20Z
The new skeptical post repeats the unresolved comparison with small local models without supplying a matched test or technical evidence about RLCD. It neither disproves Jev’s advantages nor strengthens them; the case remains a plausible specialized backend awaiting inspectable accuracy, calibration and end-to-end economics.
2026-09-17T19:21:49Z
evidence attached: reddit.post.1wj32kf — The skeptical discussion materially contextualizes the open Jev case by questioning whether its low-cost structured decisions are meaningfully beyond small local models and simple confidence estimation.
2026-09-17T18:35:03Z
Sokit adds a developer-reported tool-workflow harness lead, but no inspected code or execution results establish improved usability or performance beyond the integration activity already known. The code-review benchmark is title-only evidence: it supplies neither results nor methodology and does not yet validate Jev against hosted alternatives.
2026-09-17T18:22:33Z
evidence attached: hn.story.49744527 — This released Jev-focused tool harness materially contextualizes whether Jev's structured-decision model is usable for iterative tool workflows.
2026-09-17T18:22:33Z
evidence attached: hn.story.49744021 — The benchmark provides independent evaluation evidence about Jev's cost and quality relative to hosted frontier models.
2026-09-17T17:52:36Z
A third-party user now reports Jev controlling Pokémon Red through two gyms for under $2 in Jev tokens, adding a concrete but unverified agent-control outcome rather than another integration headline; the separately built harness and its costs prevent attributing end-to-end performance or economics to Jev. The title-only “Jevmlx” submission establishes neither an open Jev release nor a first-party implementation.
2026-09-17T17:26:01Z
evidence attached: hn.story.49743108 — The Jevmlx repository may provide a first-party implementation artifact for the open Jev structured-decision model case.
2026-09-17T17:26:01Z
evidence attached: reddit.post.1wixmux — A usable public demonstration gives concrete adoption evidence for Jev's structured-decision model and its low-cost agent-control potential.
2026-09-17T16:30:03Z
“Jev for Home Assistant” adds an integration lead, not verified outside deployment: the supplied evidence is only a title, and a Home Assistant demo was already mentioned in the launch discussion. It does not establish independent adoption or improve the evidence for accuracy, calibration or comparative economics.
2026-09-17T16:22:55Z
evidence attached: hn.story.49741371 — A third-party Home Assistant integration of the Jev model, if that is what HA-Jev is, would be the first concrete outside deployment evidence for that case's adoption question.
2026-09-17T13:41:11Z
The newly attached headline repeats Jev’s release positioning without supplying independent testing or inspectable reporting; it adds coverage, not corroboration of the performance claims. Jev remains a plausible specialized decision backend, but the case still turns on matched-workload evaluation rather than further amplification.
2026-09-17T13:22:27Z
evidence attached: hn.story.49740121 — Independent secondary coverage points to the same Jev structured-decision release, modestly corroborating its positioning as a low-latency LLM alternative.
2026-09-17T11:26:18Z
A new early-access user reports strong harmful-content grading results across four public benchmarks, adding a guardrail-evaluation lead but no named benchmarks, scores or inspectable comparisons. This modestly broadens reported usage without validating calibration or changing the performance assessment; repeated prior-art discussion adds no verified architectural evidence.
2026-09-17T11:21:45Z
evidence attached: reddit.post.1wiq7vn — A user report provides early usage and benchmark evidence consistent with Jev's claimed low-cost decision-engine positioning, though not independent validation.
2026-09-17T07:29:23Z
The new Open-jev submission advertises one-pass option scoring with Gemma 3 4B, adding another local-baseline lead but no inspectable implementation or comparative results in the supplied evidence. It does not establish equivalence to Jev or change the existing uncertainty around performance and calibration.
2026-09-17T07:22:35Z
evidence attached: hn.story.49737236 — Open-jev is a direct independent implementation of one-pass option scoring with a small local model, materially contextualizing the Jev structured-decision approach.
2026-09-17T06:22:20Z
A developer reports obtaining early access and publishes a concrete MCP connector repository for experimenting with Jev through Claude and Codex, adding a practical integration path rather than another architectural analogy. This modest ecosystem development does not validate the connector's operation or strengthen Jev's accuracy, calibration or performance claims.
2026-09-17T06:21:55Z
evidence attached: reddit.post.1wilbnz — A released MCP connector is concrete ecosystem evidence that developers are beginning to integrate Jev with established agent harnesses.
2026-09-17T05:28:42Z
The new HN submission repeats the same author's prior-art claim without identifying an inspectable paper, model, dataset or package in the supplied excerpt. Cross-platform repetition adds neither independent corroboration nor a credible contradiction of Jev's novelty or performance claims.
2026-09-17T05:21:35Z
evidence attached: hn.story.49736660 — The author points to an earlier paper, model, dataset, and package that materially contextualize Jev's claimed structured-decision architecture and reproducibility.
2026-09-17T04:28:26Z
The latest attachment is a third repetition of the same author's prior-art claim, not an additional independent implementation or an inspectable public artifact. Neither it nor the new comments changes the evidence for Jev's performance, calibration or novelty; attention remains mostly repetitive amplification.
2026-09-17T04:21:32Z
evidence attached: reddit.post.1wijo3e — Independent prior implementation and public artifact provide useful corroborating context for Jev’s claimed non-autoregressive structured-decision approach.
2026-09-17T03:25:54Z
The browser-agent submission introduces a potentially useful application lead, but its title alone does not establish a working integration or successful transfer beyond classification. The prior-art post repeats the same author's earlier claim without identifiable artifacts in the supplied evidence; neither addition materially strengthens or contradicts Jev's performance claims.
2026-09-17T03:22:11Z
evidence attached: hn.story.49735979 — This is an independent artifact using Jev's dynamic indexed action space for browser agents, materially testing whether the model's structured-decision approach transfers beyond classification.
2026-09-17T03:22:11Z
evidence attached: reddit.post.1wihgum — The author claims a prior open implementation of the same non-autoregressive structured-decision approach, adding provenance and novelty context to the Jev episode.
2026-09-17T02:33:41Z
The new prior-art claim and quoted Qwen implementation claim add leads for local comparisons, not verified evidence about Jev’s architecture, novelty or performance. The attachment rationale overstates the supplied evidence: no paper, model artifact or benchmark has been inspected, so the central assessment remains unchanged.
2026-09-17T02:22:29Z
evidence attached: reddit.post.1wih6k8 — Independent first-party artifact claims a similar non-autoregressive structured-decision architecture, materially contextualizing Jev’s novelty and provenance.
2026-09-17T00:28:15Z
The new DiffusionGemma/vLLM submission suggests another local structured-decision approach, but its title alone does not establish an implementation, upstream vLLM support, or measured performance. Implementation interest is broadening without materially strengthening Jev’s frontier-comparability or calibration claims.
2026-09-17T00:22:12Z
evidence attached: hn.story.49734375 — A vLLM implementation of Jev-like diffusion inference provides concrete serving evidence relevant to the emerging non-autoregressive structured-decision path.
2026-09-16T19:31:03Z
The new HN submission adds a claimed Jev-like reproduction and a Doom-demo link, but no inspected implementation or measured results; its attachment rationale overstates the available artifact-level evidence. This extends implementation interest without establishing independent replication, calibration or competitive performance, so the case remains cool.
2026-09-16T19:22:34Z
evidence attached: hn.story.49731282 — A reverse-engineered Jev-like implementation provides independent artifact-level evidence that the structured-decision model approach is reproducible, though not that its performance claims hold.
2026-09-16T15:40:31Z
The new OpenJev post repeats the same developer’s already-counted implementation claim, adding an option-letter/logit recipe but no measured comparison or calibration evidence. The other discussion reiterates that schema conformity does not prevent wrong decisions; neither attachment materially strengthens the case, so attention cools pending substantive testing.
2026-09-16T15:22:38Z
evidence attached: reddit.post.1whzy7j — The released openjev implementation is independent practical evidence that ordinary local model logits may approximate Jev-style structured decisions.
2026-09-16T15:22:38Z
evidence attached: reddit.post.1whzehf — shared external link with case evidence
2026-09-16T14:37:25Z
The local-model discussion adds a concrete claimed implementation, OpenJev, using Qwen 4B logits, making a local comparison more actionable than the earlier GLiClass headline. This is evidence of implementation interest, not a verified reproduction of Jev’s architecture, calibration or performance, so it does not yet warrant promotion.
2026-09-16T14:22:54Z
evidence attached: reddit.post.1whxf90 — The discussion directly probes whether Jev's structured-decision approach can be reproduced locally, materially informing the open implementation and deployment hypothesis.
2026-09-16T11:27:16Z
A reported third-party event-validation test moves Jev beyond vendor-only performance claims, providing narrow corroboration of practical speed, cost and accuracy benefits. Only the tester’s headline is available, so this strengthens the case for a matched-workload trial without validating frontier equivalence or probability calibration.
2026-09-16T11:22:12Z
evidence attached: reddit.post.1whtlzq — This is independent third-party testing of Jev reporting faster, cheaper, and more accurate event validation, directly bearing on the open product-performance hypothesis.
2026-09-16T06:28:02Z
The new attachments show wider discussion, not independent validation of Jev’s performance; the GLiClass headline alone does not establish a working alternative or replication. The opportunity remains specialized decision inference, with schema guarantees distinct from factual correctness.
2026-09-16T06:22:03Z
evidence attached: reddit.post.1whop6b — The linked Register report independently corroborates JEV's release and its decision-oriented positioning.
2026-09-16T06:22:03Z
evidence attached: hn.story.49722200 — The independent coverage corroborates the released JEV decision-model episode, though its performance claims remain unverified.
2026-09-15T21:36:24Z
grounded: converges/medium — TypeSafe AI’s software-facing decision-model direction converges with Scott’s Micro-Judgement Pattern and structured-extraction implementations: Jev could warra
2026-09-15T21:31:16Z
case created — An accessible first-party announcement, evaluation artifacts, and substantial discussion establish a moving release episode whose claims are narrower than general frontier-model equivalence.