InternLM β the lab behind the Intern-S2/InternVL model families, per the case's own assessment (the supplied snippets themselves treat it only as a Hugging Face org) β quietly uploaded three 'structured decision models', Intern-Decision-0.8B/2B/4B, to Hugging Face within roughly forty seconds on 26 September 2026, with no announcement, paper, or repository. Each is an Apache-2.0 (plus LICENSE-QWEN) fine-tune of Qwen3.5 that accepts a shared state, a schema of named questions, and up to eight images, then in a single forward pass β the card states the path 'does not call generate() or sample free-form text' β maps each field's options to single-token symbols, reads the logits immediately before placeholder positions in a rendered JSON skeleton, softmaxes over only that field's allowed candidates, applies fitted-temperature calibration, and returns Jev-compatible JSON with calibrated distributions. All performance figures are lab-self-reported: 44.16 ms mean per query on an RTX 4090 for the 4B versus Jev 1.13's 109.7 ms, and 90.02 average across seven evaluation sets versus Jev's 88.74 (Brier 0.347, ECE 0.065) β third-party coverage explicitly calls the evidence 'a vendor table unreproduced', and no independent evaluation exists. The release deliberately benchmarks into an existing small niche, naming TypeSafe's Jev and Convai's Laya as comparators, and positions one-pass decision models as drop-in calibrated routing components for agent orchestration.
| source | object | author | score | comments |
| π reddit | internlm/Intern-Decision 4B and 0.8B LocalLLaMA Retrieved article excerptOpen article Β· Retrieved 2026-09-26T09:23:39.538198+00:00 # [internlm](https://huggingface.co/internlm) / [Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) Like 1 Follow Intern Large Models 1.3k
[Image-Text-to-Text](https://huggingface.co/models?pipeline_tag=image-text-to-text)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)[qwen3\_5](https://huggingface.co/models?other=qwen3_5)[decision-making](https://huggingface.co/models?other=decision-making)[multimodal](https://huggingface.co/models?other=multimodal)[structured-prediction](https://huggingface.co/models?other=structured-prediction)[conversational](https://huggingface.co/models?other=conversational)
License: apache-2.0
[Model card](https://huggingface.co/internlm/Intern-Decision-4B) [Files Files and versions
xet](https://huggingface.co/internlm/Intern-Decision-4B/tree/main) [Community](https://huggingface.co/internlm/Intern-Decision-4B/discussions)
Deploy
Copy to bucket new
Use this model
# Intern-Decision-4B
[Demo](https://huggingface.co/spaces/internlm/intern-decision) | [Model Weights](https://huggingface.co/collections/internlm/intern-decision) | [GitHub](https://github.com/internlm/Intern-Decision)
**Intern-Decision-4B** is a multimodal structured decision model fine-tuned from
**[Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)**.
It accepts a shared state, a schema of named questions, and optional images,
and returns an answer distribution for every question in one model forward pass.
## How inference works
1. Preserve the question and option order, and map each question's options to
single-token symbols `A`, `B`, β¦, `Z`, `a`, β¦, `z`, `0`, β¦, `9`.
2. Render the original system prompt, state, decision schema, and a complete
assistant JSON skeleton with one `<decision>` placeholder per field. Preserve
the checkpoint's chat template and empty thinking block.
3. Run one causal Hugging Face forward pass. For the masked-next-token decision
objective, read logits at the position **immediately before each placeholder**.
4. Take a softmax over only that field's allowed candidate-symbol logits, then
apply the checkpoint's probability calibration.
5. Map symbols back to the original option values and return typed JSON answers.
This API performs structured candidate scoring. It does not call `generate()` or
sample free-form text. A request can contain multiple fields; no gold answers are
inserted into the prompt. The inference compiler uses only `state`, `questions`,
and optional `images`.
## Benchmark results
| Model | Jevbench-Easy | Jevbench-Original | Jevbench-Hard | Typed Decision | ToolACE | AG News | WildJailBreak | Average | Brier β | ECE β |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Jev | 100.00 | 98.61 | 72.07 | 73.35 | 91.29 | 89.57 | 96.29 | 88.74 | 0.358 | 0.095 |
| Laya | 95.83 | 72.22 | 28.83 | 35.95 | 63.87 | 92.84 | 14.84 | 57.77 | 0.804 | 0.246 |
| SemIf | 100.00 | 98.61 | 61.26 | 62.80 | 85.16 | 89.22 | 92.53 | 84.23 | 0.498 | 0.112 |
| Kev | 100.00 | 93.06 | 45.05 | 65.60 | 87.42 | 89.82 | 75.97 | 79.56 | 0.738 | 0.262 |
| JevK5 | 100.00 | 97.22 | 73.87 | 64.50 | 80.97 | 89.13 | 90.45 | 85.16 | 0.366 | 0.047 |
| Intern-Decision-0.8B | 97.92 | 80.56 | 52.25 | 77.35 | 94.52 | 88.61 | 64.48 | 79.38 | 0.530 | 0.066 |
| Intern-Decision-2B | 100.00 | 84.72 | 63.96 | 79.35 | 96.45 | 89.96 | 78.33 | 84.68 | 0.437 | 0.100 |
| Intern-Decision-4B | 100.00 | 98.61 | 73.87 | 80.55 | 96.45 | 90.82 | 89.86 | 90.02 | 0.347 | 0.065 |
## Inference latency
Measured on a single RTX 4090 with the local HF inference path. Values are
per-query end-to-end latency; they are workload and hardware dependent.
| Model | Mean | Median / P50 | P95 |
| --- | --- | --- | --- |
| Jev | 109.70 ms | 106.30 ms | 146.70 ms |
| Intern-Decision-0.8B | 33.98 ms | 33.44 ms | 37.50 ms |
| Intern-Decision-2B | 33.28 ms | 33.15 ms | 33.55 ms |
| Intern-Decision-4B | 44.16 ms | 44.03 ms | 44.60 ms |
## Known-distribution calibration pilot
This separate 96-case diagnostic uses exact reference distributions rather than
sampled hard labels. Lower is better. The pilot was not used to fit or select
the published temperature; the 4B model used its separately fitted T=1.992418.
| Category | Intern-Decision-4B before | Intern-Decision-4B after | Jev |
| --- | --- | --- | --- |
| Direct randomness and support | 0.483 / 0.181 | 0.421 / 0.129 | 0.490 / 0.216 |
| Composed events and mixtures | 0.677 / 0.254 | 0.577 / 0.150 | 0.682 / 0.274 |
| History, conditioning, and hidden state | 0.711 / 0.219 | 0.613 / 0.108 | 0.657 / 0.113 |
| Daily evidence and observation bias | 0.701 / 0.328 | 0.575 / 0.210 | 0.603 / 0.114 |
| Selective disclosure and probability puzzles | 0.540 / 0.119 | 0.510 / 0.049 | 0.483 / 0.138 |
| Sequential and combinatorial processes | 0.656 / 0.180 | 0.605 / 0.058 | 0.657 / 0.116 |
| **Overall (Brier / ECE)** | **0.628 / 0.213** | **0.550 / 0.089** | **0.595 / 0.130** |
## Quick start
Use **Python 3.12+**. Install `requirements.txt` in a suitable PyTorch/CUDA
environment, then import `DecisionEngine` from the downloaded model directory:
```
pip install -r requirements.txt
```
```
from inference import DecisionEngine
engine = DecisionEngine(device="cuda") # Load once; reuse for subsequent requests.
request = {
"state": "The customer was charged twice and asks for the extra payment back.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments and refunds",
"delivery": "Shipping and delivery",
},
},
"urgency": {
"type": "score",
"instructions": "Rate the priority.",
"criteria": ["Low", "Medium", "High"],
},
"refund_requested": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
},
}
response = engine.predict(request) # One Python dict in, one response dict out.
print(response["answers"])
```
`predict(request)` accepts one request dictionary per call and returns a
JSON-serializable Jev-compatible response. It does not read request files or mutate
the supplied dictionary. Reuse the engine for each subsequent request.
The engine defaults to the checkpoint next to `inference.py`. To load another
local copy of this same model, use `DecisionEngine(checkpoint="./model-copy")`.
Use the inference module shipped with the selected size so its default calibration
matches. `backend="hf"` is the default and the only implemented backend. The
optional request `model` field does not switch checkpoints; the response `model`
identifies the weights actually loaded by this module.
### Request format
```
{
"state": "The customer was charged twice and asks for the extra payment back.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments and refunds",
"delivery": "Shipping and delivery"
}
},
"urgency": {
"type": "score",
"instructions": "Rate the priority.",
"criteria": ["Low", "Medium", "High"]
},
"refund_requested": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}
```
- **choice**: `criteria` is an ordered object mapping option values to descriptions.
- **score**: `criteria` is a list (values become `"0"`, `"1"`, β¦) or an ordered
object with finite numeric string keys.
- **noul**: a binary decision with options `no`, then `yes`. Optional criteria can
describe these values using `no`/`yes` or `false`/`true` keys.
Supply 1β16 questions, with up to 62 options per question. Inputs exceeding
`DecisionEngine(max_length=8192)` (default 8192 tokens) are rejected without truncation.
### Images
Set the request dictionary's `images` list in the intended order:
```
request["images"] = ["images/frame-1.png", "images/frame-2.png"]
response = engine.predict(request)
```
The checkpoint processor handles image resizing and token expansion. Relative
paths are resolved against `DecisionEngine(media_root=".")` (default: the working directory).
Supply up to eight images; image tokens count toward the input length limit.
### Response format
`answers` maps each field name to:
| Field | Meaning |
| --- | --- |
| `type` | `choice`, `score`, or `noul` |
| `probabilities` | Calibrated distribution over the original option values |
| `confidence` | Maximum candidate probability |
| `decision` | Highest-probability option value; lexical tie-breaking |
| `choice` | Selected value, for choice questions |
| `noul` | Probability of `yes`, for binary questions |
| `score` | Probability-weighted expected numeric value, for score questions |
| `legend` | Score values and their descriptions, for score questions |
| `source` | `local` |
The response follows the Jev envelope: `model`, `answers`, and `usage`.
It also includes `backend`, `timing`, and `calibration` as extension fields.
`usage.output_tokens` and `usage.decision_count` count scored fields, not generated
text tokens. `confidence` for a score question belongs to its most likely category;
the reported expected `score` can lie between categories.
## Calibration
The default temperature is **1.99241824**. It was fitted separately for this checkpoint
by NLL minimization on 1,728 designated calibration cases, with 1,693 separate
validation cases. Test-suite labels were not used to select the temperature.
The script follows the demo's numerical sequence:
```
p = softmax(candidate_logits.float())
calibrated_p = softmax(log(p) / T)
```
This is candidate probability calibration, **not a sampling temperature**. It
updates confidence, the `noul` probability, and the expected `score` while
preserving the argmax decision. For uncalibrated candidate probabilities, use
`DecisionEngine(temperature=1)`. A custom temperature must be finite and positive.
## License and acknowledgment
Intern-Decision is derived from the Qwen3.5 series. The original Qwen license is preserved
as [LICENSE-QWEN](https://huggingface.co/internlm/Intern-Decision-4B/tree/main/LICENSE-QWEN). Retain the license and applicable
upstream notices when redistributing. These weights were modified by decision
tuning, and this release adds the structured inference wrapper and model card.
We thank the Qwen team for the original models and multimodal processor.
Downloads last month
: -
Safetensors
Model size
5B params
Tensor type
F32
Β·
BF16
Β·
Chat template
Files info
## Model tree for internlm/Intern-Decision-4B
Base model
[Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base)
Finetuned
[Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
Finetuned
([784](https://huggingface.co/models?other=base_model:finetune:Qwen/Qwen3.5-4B))
this model | jacek2023 | 41 | 10 |
| π§ echo.github β | Model card: Intern-Decision 4B/0.8B are multimodal structured decision models fine-tuned from Qwen3.5 that 'return an answer distribution fo | InternLM (internlm) | β | β |
| π§ hn | Julia-1: decision model that runs on almost anything | handfuloflight | 4 | 0 |
| π§ hn | Show HN: Peekaboolean β image and jev-like typed questions in typed anwers out | bykof | 2 | 0 |
| π reddit | I trained a 500M VLM that answers typed questions about an image (choice / score / yes-no) with calibrated probabilities. ~400 ms on an M1 Pro, no text generation [P] MachineLearning | bykof | 1 | 5 |
| π reddit | Better, Faster, and More Calibrated than Jev, with Multi-Modal Ability [P] MachineLearning | Spico197 | 1 | 1 |
| π§ hn | Run Decision Models on vLLM and Red Hat AI Using DiffusionGemma | thebeardisred | 1 | 0 |
| π§ hn | d1: Liquid AI's First Decision Model | mfiguiere | 3 | 0 |
| π§ hn | Liquid AI releases decision model D1 | mnewme | 1 | 0 |
| π reddit | stuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model LocalLLaMA | Inevitable-Log5414 | 1 | 0 |
| π reddit | Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38Γ faster decisions and +8.7 points accuracy, for under 2 GB extra memory LocalLLaMA | Usual_Maximum7673 | 114 | 32 |
| π§ hn | Free hosted API for Laya, the open-weight decision model | boundlesshq | 1 | 1 |
| π reddit | building an open source coding agent called Z-Engine OpenAI | arshadbarves | 0 | 2 |
| π§ hn | New in Llama.cpp: Decision Models | cuvinny | 6 | 1 |
| π reddit | DecisionTune 1.0: a 395M encoder that picks from your options offline, about 10 ms per short decision on MLX (Apache-2.0) LocalLLaMA | Abe238 | 14 | 1 |
| π reddit | ARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti LocalLLaMA | KMatysek | 2 | 2 |
2026-10-06T14:44:56Z
ARC-1 (1.7B LoRA, order-invariant per-option branch scoring, ~20ms on a 4060 Ti, own JevBench/DecideBench numbers) is a ninth independent line β periphery keeps expanding into a long tail of small local builds while the attention arc stays dead (~0.33 pts/h vs ~73 peak). Case meaning is unchanged: the category consolidates as quiet infrastructure; heat stays low on attention while state holds accelerating on implementation spread, and every decisive trigger (independent Intern-Decision eval, d1's mechanism, model-level serving beyond backend='hf') remains open.
2026-10-06T13:34:27Z
evidence attached: reddit.post.1wz14pv β Independent 1.7B one-pass decision model with published JevBench/DecideBench numbers corroborates one-pass decision components as a spreading pattern.
2026-10-05T09:32:57Z
grounded: converges/high β Converges at Scott's strongest seam: a credible lab shipping Apache-2.0 Qwen3.5 fine-tunes that answer a named-question schema with calibrated per-field distrib
2026-10-05T09:24:50Z
DecisionTune 1.0 (395M ModernBERT encoder + scoring head, ~10 ms on MLX, Apache-2.0) is the first pure-encoder and first MLX-native entrant: the one-pass typed-decision pattern now spans decoder fine-tunes, distilled proxy heads, LoRA-routed routers and encoder scorers across CUDA/llama.cpp/MLX β architectural convergence on the pattern, not more of the same build. Scott-facing meaning is unchanged: Intern-Decision's figures remain lab-self-reported and the decisive triggers (independent eval, d1 mechanism, serving beyond backend='hf') stay open.
2026-10-05T09:23:01Z
evidence attached: reddit.post.1wy3c0u β Independent parallel implementation of the one-pass typed-decision pattern (395M encoder, ~10ms local routing) corroborates the accelerating hypothesis that small decision models become drop-in agent routing components.
2026-10-02T15:45:34Z
This pass fully prices the llama.cpp development attached last cycle: decision-model support is now native in the canonical local runtime (ggml-org shipping it; comments surface gutsy 0.8B, another CPU-capable llama.cpp-compatible indie), moving the category's deployment story from HF-backend-only toward runtime-native infrastructure β though Intern-Decision itself still serves only via backend='hf' and every figure remains lab-self-reported. Jeff's thread shows only a mild tail re-spike (~3 pts/h, 3.6Γ a decayed baseline), so the measured 'accelerating' momentum reflects decay arithmetic, not a new wave β heat holds low.
2026-10-02T15:25:28Z
evidence attached: hn.story.49934323 β ggml-org shipping decision-model support in llama.cpp is direct ecosystem-adoption evidence for one-pass decision models becoming drop-in routing components.
2026-10-02T03:40:56Z
First adopter-side evidence β the Z-Engine open-source coding agent embedding Laya for approval/routing/escalation micro-decisions β marks the pattern crossing from model-builders to agent-consumers, a new kind of periphery line even though its traction is thin. Meanwhile the attention wave that justified medium has crested and decayed: Jeff's single thread fell from ~55 pts/h peak to ~4.5 with cooling momentum, so the spike is confirmed as single-threaded and passing. Heat cools mediumβlow as an attention judgment, not a belief change β the category still ticks over (new adopters appearing) and Intern-Decision's figures remain lab-self-reported with every named trigger (independent eval, d1 mechanism, serving beyond backend='hf') still open, so state holds at accelerating on breadth.
2026-10-02T03:26:03Z
evidence attached: reddit.post.1wvibcd β Independent builder adopting a specialized model for agent micro-decisions (approval, routing, escalation) β independent corroboration of the one-pass routing-component pattern.
2026-10-01T22:12:25Z
The velocity spike is real but single-threaded: Jeff's LoRA-router post has become the case's attention engine (9β64 pts/24 comments, 5Γ baseline, 90th peer percentile, accelerating) and now carries the first methodological critique β 300-row slices without CIs, post-quantization calibration drift, coverage-vs-accuracy abstention β which marks maturing category discourse, while Laya's free hosted API on Vercel's gateway shows decision models beginning to ship as served utilities rather than weights-only releases. Neither changes claim status β Intern-Decision's figures remain entirely lab-self-reported with no independent eval, d1's mechanism is still unknown, and serving beyond backend='hf' is still absent β so the case holds at accelerating with heat lifted lowβmedium because attention is live and the periphery is still expanding, not because belief moved.
2026-10-01T20:35:48Z
evidence attached: hn.story.49926327 β Independent corroboration: a second open-weight decision model (Laya) now freely hosted on Vercel's gateway is spread evidence for one-pass decision models as agent-routing components.
2026-10-01T15:16:00Z
A fifth independent implementation (Jeff-Qwen3.5-0.8B v1.2: calibrated one-pass 0.8B classifier with 9 hot-swapped per-task LoRA adapters in front of a 27B, confidence-gap-based forwarding, claimed 38Γ faster/+8.7 pts at <2GB extra) shows the indie envelope still broadening β from bare one-pass classifiers to composed router stacks β and draws the case's first substantive methodological critique (300-row slices without CIs, post-quantization calibration drift, coverage-vs-accuracy abstention curves), so the category discourse is maturing even though headline attention stays modest: the 1.5 pts/h / 68th-percentile tick-up is entirely Jeff's own traction, not platform reignition. Intern-Decision's figures remain single-sourced and every named trigger (independent eval, d1's mechanism, serving beyond backend='hf') is still open, so the case holds at accelerating on breadth, not attention.
2026-10-01T14:32:26Z
evidence attached: reddit.post.1wv05u1 β Independent parallel implementation of the one-pass decision-model pattern (0.8B classifier with per-job LoRAs routed in front of a 27B) corroborating that specialized decision models are emerging as agent components.
2026-09-30T20:19:12Z
stuntd 0.1.2 makes the periphery expand again after the last look judged it stopped: a fourth indie implementation, and architecturally new β a proxy that distils an LLM's own typed decisions into small local heads (~20ms GPU/~60ms CPU) with multi-field all-or-nothing confidence gating, i.e. the distill-your-own-decision-logs route rather than a shipped decision model. The category thesis thus broadens from 'trained decision models' to 'distilled routing proxies' even while raw attention sits at zero, so the case holds accelerating at low heat on its unchanged triggers: independent Intern-Decision eval, d1's actual mechanism, serving/framework support beyond backend='hf'.
2026-09-30T18:42:39Z
evidence attached: reddit.post.1wu8x61 β stuntd's local multi-field decision heads with all-or-nothing confidence gating is independent spread of the one-pass decision-model-as-routing-component pattern.
2026-09-30T06:37:50Z
The only new evidence is hn.story.49904832, a duplicate HN submission (1pt/0 comments, title-level) of the Liquid AI d1 story already priced β no new fact, no mechanism detail, no independent eval β so the case's meaning is unchanged and the alert-delivery 'high' settles back to low heat now the heads-up has been delivered. Attention is fully decayed (0.17 pts/h at 94h, 40th percentile, zero comments/h) while the periphery has stopped expanding since the d1 attach, so the case holds at accelerating/low and waits on its named triggers: independent Intern-Decision eval in others' hands, what d1's mechanism actually is, and serving/framework support beyond backend='hf'.
2026-09-30T06:23:29Z
evidence attached: hn.story.49904832 β shared external link with case evidence
2026-09-29T22:29:03Z
grounded: converges/high β Converges at the strongest seam available: a credible lab has shipped, as Apache-2.0 Qwen3.5 fine-tunes, the trained form of Scott's micro-judgement/nudge-doctr
2026-09-29T22:22:07Z
relevance=high case never alerted; deterministic escalation to deliver
2026-09-29T20:53:12Z
evidence attached: hn.story.49899623 β Liquid AI shipping its own decision model is independent corroboration that one-pass decision models are forming a real vendor category.
2026-09-29T04:38:12Z
The category acquires an infrastructure dimension: decision-model serving is documented on mainstream vLLM/Red Hat AI stacks via DiffusionGemma (title-level evidence, 1pt/0 comments, content unverified, and a different model than Intern-Decision), partially answering the previously open 'will mainstream serving stacks accommodate one-pass decision models?' question and bearing on any drop-in-routing deployment path. It does nothing for claim-level verification β every Intern-Decision figure remains lab-self-reported and its only implemented backend is still the local HF forward pass β so the case holds at corroborated with low heat, the serving-stack path now flagged as the concrete trigger to watch for acceleration.
2026-09-29T04:26:18Z
evidence attached: hn.story.49887652 β Decision-model serving documented on mainstream vLLM/Red Hat infrastructure materially contextualises whether one-pass decision models become standard orchestration routing components.
2026-09-28T22:29:08Z
The two newly attached posts change nothing material: reddit.post.1wrgbhi is the InternLM team's own Reddit promotion of the already-verified 0.8B/2B/4B family and its Jev 1.13.0 comparison β no claims beyond the model card we fetched β and reddit.post.1wrf4sm merely crossposts Peekaboolean (already priced as the third independent implementation) to Reddit at 0 pts. Meaning is unchanged: category-level corroboration stands, every Intern-Decision calibration/latency figure remains lab-self-reported, and engagement is fully decayed (0.0 pts/h at 61h, main post slipping 43β40) so heat stays low despite the hot agent-orchestration neighbourhood.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wrgbhi β Same Intern-Decision project expanding to a multi-modal 0.8B/2B/4B family with a claimed Jev-beating 4B and ~30 FPS on a 4090, directly extending the one-pass decision-model case.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wrf4sm β shared external link with case evidence
2026-09-27T11:46:19Z
Peekaboolean β an independent VLM answering choice/score/binary typed questions about images at ~400ms p95 on an M1 Pro, in the same 'Jev' vocabulary β is the third independent implementation of the one-pass typed-decision pattern after IdeaNJEV and Julia-1, and that periphery (plus the Jev-lineage set Intern-Decision itself benchmarks against) crosses the corroboration bar for the pattern thesis: one-pass typed-decision routing components are now an independently reproduced category, not a single-lab drop. What remains uncorroborated is Intern-Decision itself β every calibration/latency/benchmark figure is still lab-self-reported with no independent eval, framework adoption, or populated download count β so this is category maturity, not claim verification; engagement is thin and quiet (0.17 pts/h, 38th percentile, long decayed from the 16.8 spike) and heat stays low despite the hot agent-orchestration neighbourhood.
2026-09-27T11:24:10Z
evidence attached: hn.story.49865111 β Independent implementation extends the one-pass typed-decision pattern to vision inputs at comparable latency β pattern-spread evidence a re-judge should weigh.
2026-09-26T22:54:02Z
The new HN post (Julia-1, 3 pts / 0 comments) adds a third thin entrant to the one-pass decision-model niche, mildly reinforcing the category-forming reading already priced via IdeaNJEV and the Jev-lineage comparison set β but it concerns a different model and neither verifies nor challenges Intern-Decision's self-reported numbers. Meaning is unchanged: a real, fully documented release whose calibration/latency claims still await independent evals or framework adoption; engagement is cooling (1.2 pts/h, 65th percentile, momentum cooling), so heat stays low.
2026-09-26T22:24:09Z
evidence attached: hn.story.49860793 β A third entrant claiming a portable one-pass decision model is category-forming evidence for decision models as routing components, though the post itself is thin and uncorroborated.
2026-09-26T20:52:45Z
The velocity-spike flag decayed into a +1-upvote, zero-new-comment plateau β transient amplification, not substance; the spike never represented new evidence and doesn't move heat. State settles seedβwatching on the strength of the already-grounded, fully documented release (full model card, weights, repo, demo), not on attention: the artifact is real, its calibration/latency claims remain lab-self-reported, and the case now waits on independent evals or adoption rather than signal.
2026-09-26T09:40:39Z
grounded: converges/high β A credible lab ships, as released Qwen3.5 fine-tune weights, the trained form of what Scott's canon holds as composition doctrine: micro-judgement/Decision-DAG
2026-09-26T09:33:16Z
case created β Credible lab's released artifact establishing a distinct pattern (one-forward-pass structured decision models) not covered by any open case's specific claim.