OpenAI announced the Decisions API at DevDay on September 29, 2026: a limited-preview endpoint powered by a specialized build of GPT-6 Luna, in which a developer defines a question and a finite set of predefined answers, supplies text or image context, and gets one typed answer back for classifying content, routing requests, or choosing an agent's next action β launch coverage reports ~150ms decisions and The New Stack frames it explicitly as OpenAI's answer to TypeSafe's Jev decision models. The announcement carried no pricing, accuracy, or response-time figures, and access stayed locked to selected customers (standard keys still returning 403 'not enabled', no docs page, in early October); the case's own evidence trail then records the preview reaching public beta on October 7, on which the supplied web snippets are thin, with a first independent eval (sub-600 calls via OpenRouter against Jev and Mercury Decide) landed but its verdict not visible. Still unestablished in the excerpts in hand: official pricing (only an unofficial Luna-rate planning figure circulates), whether the API is meaningfully different from plain structured outputs on Luna, and what one article's '28.7% Luna footnote' actually refers to.
OpenAI β the most consequential party available β has moved a Luna-backed typed-decision API from stall-scare to public beta (verified by third-party usage, not official docs), independently productizing the exact cheap-model front door Scott already runs through Jev in three projects; with the first independent eval (Topfi via OpenRouter vs Jev and Mercury Decide) landed and pricing-vs-Jev still unanswered, this is the decisive window: benchmark Luna-backed decisions against Jev/Clef/Mercury Decide now, keep the front door provider-agnostic behind task-aware routing per his model-barbell and model-perishability doctrine, and hold Clef as the image-input fallback. It is simultaneously the dated-receipts publishing moment β the category leader arriving at answers-not-content as a first-class API validates his decisions-as-primitive position, and the outcome will decide whether a live dependency in Venture World, dev-wiki and all_in_one_software holds or gets swapped.
dev:technology.typesafe-jevdev:project.jevdev:concept.cheap-model-front-doorip:concept.model-barbelldev:concept.task-aware-model-routingip:concept.model-perishabilityradar:typesafe-jev-structured-decisionsradar:person.typesafe-airadar:concept.typed-decisionsradar:concept.structured-decisions
queries asked of Scott's wikis
- typed decision interface primitive in production agent stacks
- Jev / TypeSafe dependency across dev projects β where and how load-bearing
- provider-agnostic task-aware routing layer for model calls
- specialist decision models vs general LLM structured outputs β latency and cost positions
- commodity argument for decision models β roll-your-own vs specialist API
- governor module / decision governance in agent orchestration writing
2026-10-08T21:03:48Z
New tutorial (hn.story.50007041) showing Jev + Decisions API integration adds another thin adoption data point from the same HN/Vercel ecosystem, but the decisive fork β production standardization on OpenAI's first-party API vs Jev-class specialist economics β still waits on the same three missing items: official pricing, Topfi's ~600-call eval verdict via OpenRouter, and any rigorous benchmark comparing the Decisions API (not raw Luna) against Jev/Clef/Mercury Decide. Engagement continues its cooling tail (0.17 pts/h vs 87 peak). Case remains corroborated on a quiet watch.
2026-10-08T17:49:20Z
evidence attached: hn.story.50007041 β Tutorial demonstrating Jev integration with OpenAI Decisions API, evidence of early adoption patterns for the structured-decision interface.
2026-10-08T03:48:37Z
The Vercel article (hn.story.50001449) adds another thin adoption data point from the same ecosystem (flashbrew, HN), but the decisive fork still waits on official pricing, Topfi's verdict, and a rigorous benchmark. Engagement continues its cooling tail (0.17 pts/h vs 81 peak). Case remains corroborated on a quiet watch.
2026-10-08T03:35:14Z
evidence attached: hn.story.50001449 β Vercel article detailing practical use cases for OpenAI's Decisions API, bearing on adoption evidence for the structured-decision interface.
2026-10-07T19:40:59Z
The quiet watch earns its first non-amplification increment: llm-OpenAI-Decisions 0.1a0 moves Decisions API support from one-off workflows (Vercel triage, simonw's curl walkthrough) into shareable tooling in Willison's llm CLI ecosystem β thin (3pts/0cmts) but real periphery expansion on the adoption axis this case resolves on. It does not make the case accelerating: momentum is cooling (9.8/h vs 79 peak) and the decisive missing items β official pricing, Topfi's verdict, any rigorous benchmark β are all still absent, so the case stays corroborated on the same three re-heat triggers. Heat stays low despite the magnitude-valve spread flag: the 90th-percentile reading is the Oct 7 beta thread's decaying tail across the same 3 platforms, not new periphery.
2026-10-07T18:32:49Z
evidence attached: hn.story.49995246 β Independent ecosystem integration β Willison shipping llm-CLI support for the Decisions API β is exactly the early third-party adoption evidence this case resolves on.
2026-10-07T17:26:55Z
The increment since the last look is amplification only: the beta thread grew to 373pts/218cmts and a velocity flag fired on the same Oct 7 event's tail, but no new implementation, platform, community, or substantive comment appeared, and official pricing plus Topfi's verdict are still absent β so the case demotes from accelerating to corroborated (verified substance, no longer moving) and cools to a quiet watch. The deterministic spread flag and 95th-percentile rate are the beta event's decaying tail, not expanding periphery β 17.7/h is a quarter of peak, the newest substantive item (armcat's calibration probe) earned 1pt/0cmts, and chatter velocity will not produce the missing pricing or eval verdict; re-heats on official pricing, the Topfi verdict, or any rigorous benchmark. Scott's up-vote says the briefing landed; relevance stays high on his live Jev dependency.
2026-10-07T13:49:07Z
The case gains its first empirical quality signal and it cuts against the standardization branch: a thousand-trial calibration probe (armcat) reports the Decisions API's probabilities well-behaved on predicate questions but odd on choice questions β the core primitive β so quality calibration now sits beside Jev-economics as an open axis of the adopt-or-stall fork. Meanwhile the beta thread has cooled into amplification with official pricing and Topfi's verdict still missing, decaying the decisive window from hours-urgent to a daily watch that re-heats the moment either lands.
2026-10-07T13:27:20Z
evidence attached: hn.story.49992182 β Hands-on empirical probing of the Decisions API's probabilistic behavior (calibrated on predicates, odd on choice) is real usage evidence bearing on the case's adoption/quality question.
2026-10-07T08:38:45Z
The economics half of the decisive fork is no longer silence: the first in-thread pricing-vs-Jev analysis (kna1) claims Luna runs >2x Jev per-token but argues cached-input pricing β which Jev lacks β flips the math when a shared option list amortizes across queries; counteracting it, Balance- reports Decisions and text generation run separate caching/inference paths, so decide-then-generate on a large input processes it twice. Beta coverage also crossed to Reddit with real substance (26pts/8cmts) rather than the earlier 1-point duplicate; Topfi's eval verdict and official pricing are still missing, so the adopt-vs-Jev-economics fork stays open.
2026-10-07T08:25:40Z
evidence attached: reddit.post.1wzprt3 β shared external link with case evidence
2026-10-07T05:02:13Z
The increment since the last look is amplification, not substance: the beta thread cooled to steady top-decile velocity (~half its 56/h peak), the only new attachment is a 1-point duplicate of the public-beta news, and the decisive missing data β Topfi's eval verdict, pricing-vs-Jev β is still absent from visible comments. The case moves to accelerating on implementation substance already in evidence (Topfi's live-API eval via OpenRouter, a shipped Vercel PR-triage workflow) plus 3-platform spread; the adoption-vs-Jev-economics fork the case exists to resolve remains open and should land in-thread within hours.
2026-10-07T03:33:14Z
evidence attached: hn.story.49986718 β Decisions API advancing from limited preview to public beta bears directly on the case's adopt-or-stall resolution.
2026-10-07T03:33:14Z
evidence attached: hn.story.49987172 β A shipped Vercel PR-triage workflow built on the Decisions API is early third-party adoption evidence the case explicitly tracks.
2026-10-07T02:13:21Z
grounded: converges/high β OpenAI β the most consequential party available β has moved a Luna-backed typed-decision API from stall-scare to public beta (verified by third-party usage, not
2026-10-07T02:04:13Z
The stall branch closed: the Decisions API shipped to public beta (HN front page, 149pts/60 comments and climbing within the hour), and the beta's reality is independently verified by usage β a third party ran sub-600-call evals against it via OpenRouter against Jev and Mercury Decide β which passes promotion on implementation substance, not engagement. The case's fork collapses to exactly its intended decisive test (production adoption vs Jev-class specialist economics), and it moves from dead-quiet to top-decile velocity: the benchmark window Scott was waiting for is open now, with the first eval verdict and the pricing-vs-Jev answers likely landing in this thread within hours.
2026-10-06T23:36:37Z
evidence attached: hn.story.49984025 β Direct development on the case: the Decisions API advancing from limited preview to public beta is exactly the stall-or-standardize test the case tracks.
2026-10-06T22:12:14Z
The case's sole concrete demand-side datapoint flipped: the Jev user who was waiting on the Decisions API for image input reports adopting Cloudflare's Clef instead β first adoption evidence for a category competitor and confirmation that image-input is a live differentiator β while the newly attached Jev-vs-frontier-smalls comparison is title-only, zero-comment, and never tests the Decisions API, so the perf-vs-Jev question stays open. Delay is now ~8 days past 'coming days', but the negative branch is a same-platform echo plus one defection, not a confirmed stall: the fork stays open, state holds at watching, and attention is dead (~0.2 pts/h at 200h, 30th percentile) so heat stays low despite the hot openai/agent-orchestration neighbourhood.
2026-10-06T20:42:17Z
evidence attached: hn.story.49982760 β Hands-on comparison of Jev against frontier small models materially contextualises whether Jev-class specialists hold up as the typed-decision layer or get squeezed by general models.
2026-10-05T12:44:14Z
First direct evidence lands on the stall branch: a week past the promised 'coming days' broad release there is no pricing, no docs page, and standard keys get 403 'not enabled' β a single anonymous but concrete, trivially checkable Reddit report, with only unverified 'private alpha' chatter on the other side. The case's meaning shifts from 'imminent-GA trigger pending' to 'delay accumulating toward stall', but a slipped preview week at OpenAI is not a confirmed stall, so the episode stays open on its original fork (GA-with-numbers-and-adoption vs confirmed stall).
2026-10-05T12:24:46Z
evidence attached: reddit.post.1wy686i β Direct evidence on the case's stall branch: a week past the promised broad release there is no pricing, no docs, and standard keys get 403 'not enabled'.
2026-10-03T15:55:08Z
The lone on-record head-to-head inverts: commenters quoting the 'Hard-Decisions' article show its benchmark ran raw gpt-6-luna with reasoning_effort 'none' β never the Decisions API itself β so the 'Luna can't compete with Jev' claim is methodologically void and the perf-vs-Jev question reverts to fully open with no usable third-party comparison in existence. Attention is dead (0 pts/h, 12th percentile, 121h) with GA still unseen; the only remaining residue is concrete-but-unverified Clef/Kev pointers and an unverified private-alpha report.
2026-10-02T00:42:01Z
First head-to-head skepticism of Luna vs Jev surfaced (title-only, article content uncaptured), and its sole commenter β calling the comparison apples-and-oranges β named Cloudflare's Clef decision models and jaredpalmer's Kev as 'middle path' alternatives, reframing the category from a two-horse race into a possibly widening field that cuts against any single standard emerging. All of it is unverified comment-level lead, so state, heat and relevance hold; the case still turns on GA and numbers, now with Clef/Kev added to the benchmark watchlist.
2026-10-01T23:31:30Z
evidence attached: hn.story.49927854 β Direct third-party head-to-head arguing Luna/Decisions API underperforms Jev β material context for the adoption-vs-stall resolution of this case.
2026-10-01T22:32:21Z
The 'governor module' HN post (score 2, zero comments, article body not captured β pipeline attributes TechCrunch, but nothing verifiable in the object) is category context showing the typed-decision/agent-governance frame entering wider discourse, not verification of the launch's specifics or of adoption; the case's meaning is unchanged and still hinges entirely on the limited preview reaching GA (promised 'in upcoming days'). Moving seed β watching purely as a tracking posture given accumulated weak signals (verified substrate, a demand-side Jev-user hint, thin third-party coverage) and an imminent GA trigger; heat stays low at ~0.67 pts/h β the 76th peer percentile reflects a quiet cohort, not a hot episode.
2026-10-01T20:35:47Z
evidence attached: hn.story.49924516 β TechCrunch coverage of OpenAI's Jev-class decision model as an agent 'governor module' is spread/contextual evidence for whether the Decisions API becomes the standard structured-decision interface.
2026-10-01T19:45:35Z
Launch attention has decayed β HN thread drifting down to score 1, ~0.3 pts/h at 77h, no implementations, benchmarks, pricing, or independent coverage of the claim itself β so the episode now hinges entirely on whether limited preview reaches GA. The one new datum, an anonymous Jev user publicly waiting to switch for image-input capability, is a demand-side hint at the squeeze dynamic but not adoption evidence; cooling to low and holding at seed until GA, numbers, or verified third-party coverage arrive.
2026-10-01T18:31:30Z
evidence attached: reddit.post.1wv4umd β A Jev user publicly waiting to switch to the Decision API is direct demand-side evidence for the case's squeezing-dynamic resolution criterion.
2026-09-30T05:42:53Z
grounded: converges/high β OpenAI productising typed decisions as a first-party API is a consequential party independently arriving at the interface primitive Scott already runs in produc
2026-09-30T05:36:04Z
case created β First-party OpenAI launch (451K-view X announcement) that directly contests the category TypeSafe's Jev case targets β its own claim and trajectory, not merely context on an existing case.