2026-10-11 16:37 UTC

InternLM claims its released Intern-Decision 4B/0.8B models return calibrated answers to a named question schema from shared agent state in a single forward pass (~34–44 ms per query on an RTX 4090), positioning specialized one-pass decision models as drop-in routing components for agent orchestration.

state: acceleratingheat: lowuncertainty: mediumconvergesscott: highagent-orchestration local-models structured-decision-modelsInternLM
Surfaced 2026-09-29T23:22:21Z β€” Model card: Intern-Decision 4B/0.8B are multimodal structured decision models fine-tuned from Qwen3.5 that 'return an answer distribution fo β€” Liquid AI shipping d1 β€” its first decision model β€” is the category's first established-vendor entrant, joining three indie one-pass implementations (IdeaNJEV, Julia-1, Peekaboolean) and the vLLM/Red Hat serving signal; one-pass typed decision models now read as a forming vendor category rather than an indie-reproduction pattern, so the case moves corroborated β†’ accelerating on implementations plus an influential entrant. Intern-Decision's own figures remain lab-self-reported with no independent eval, and attention is fully decayed (0.17 pts/h at 85h), so heat stays low.

What is this?

InternLM β€” the lab behind the Intern-S2/InternVL model families, per the case's own assessment (the supplied snippets themselves treat it only as a Hugging Face org) β€” quietly uploaded three 'structured decision models', Intern-Decision-0.8B/2B/4B, to Hugging Face within roughly forty seconds on 26 September 2026, with no announcement, paper, or repository. Each is an Apache-2.0 (plus LICENSE-QWEN) fine-tune of Qwen3.5 that accepts a shared state, a schema of named questions, and up to eight images, then in a single forward pass β€” the card states the path 'does not call generate() or sample free-form text' β€” maps each field's options to single-token symbols, reads the logits immediately before placeholder positions in a rendered JSON skeleton, softmaxes over only that field's allowed candidates, applies fitted-temperature calibration, and returns Jev-compatible JSON with calibrated distributions. All performance figures are lab-self-reported: 44.16 ms mean per query on an RTX 4090 for the 4B versus Jev 1.13's 109.7 ms, and 90.02 average across seven evaluation sets versus Jev's 88.74 (Brier 0.347, ECE 0.065) β€” third-party coverage explicitly calls the evidence 'a vendor table unreproduced', and no independent evaluation exists. The release deliberately benchmarks into an existing small niche, naming TypeSafe's Jev and Convai's Laya as comparators, and positions one-pass decision models as drop-in calibrated routing components for agent orchestration.

Why it matters to Scott

Converges at Scott's strongest seam: a credible lab shipping Apache-2.0 Qwen3.5 fine-tunes that answer a named-question schema with calibrated per-field distributions in one forward pass is the trained, released form of his micro-judgement/Decision-DAG/nudge doctrine β€” and the surrounding periphery (eight independent builds, llama.cpp-native decision-model serving) corroborates the category, so this is a dated receipt, not the radar merely re-tracking its own typed-decisions lineage. It also poses a concrete act, not just a confirmation: Intern-Decision's self-reported Jev-beating figures (90.02 vs 88.74; ~44 ms vs ~110 ms on a 4090) bear directly on the Jev 1.13.0 front doors he has already productionised and on the gamepc/Ollama deployment path now materially eased by llama.cpp β€” with every performance claim still lab-self-reported, so the actionable trigger remains an independent eval in others' hands.
ip:framework.nudge-doctrineip:concept.micro-judgement-patternip:concept.decision-dagdev:technology.typesafe-jevdev:project.jevdev:concept.cheap-model-front-doordev:project.gamepcip:concept.transcript-distillationradar:concept.typed-decisionsradar:typesafe-jev-structured-decisionsradar:llamacpp-decision-modelsradar:verdict-local-jev-compatible-decisionsradar:system-one-lite-typed-decisionsradar:blink-embedded-typed-decisionsradar:openai-decisions-apiradar:concept.structured-decisions
queries asked of Scott's wikis
  • micro-judgement decision DAG nudge doctrine
  • calibrated confidence routing abstention gating
  • typesafe-jev front door production stack
  • distill decision logs into small router model
  • local inference estate Ollama llama.cpp deployment
  • logit scoring versus generation for structured decisions

Measured heat

now 0 pts/hpeak 56 pts/hcomments 0/hpeers p50momentum: steady3 platformsage 367h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-26 09:33 (minted)⭐ origin echo-reconstructedModel card: Intern-Decision 4B/0.8B are multimodal structured decision models fine-tuned from Qwen3.5 that 'return an answer distribution fo
InternLM (internlm) on github (echo) Β· attributed from reddit.post.1wqloas Β· published time unknown
β€”
09-26 08:54first on r/LocalLLaMA Β· published Β· lag ?internlm/Intern-Decision 4B and 0.8B
jacek2023
β€”
09-26 21:40first on hacker news Β· published Β· lag ?Julia-1: decision model that runs on almost anything
handfuloflight
β€”
09-27 08:46first on r/MachineLearning Β· published Β· lag ?I trained a 500M VLM that answers typed questions about an image (choice / score / yes-no) with calibrated probabilities. ~400 ms on an M1 Pro, no text generation [P]
bykof
β€”
10-02 02:28first on r/OpenAI Β· published Β· lag ?building an open source coding agent called Z-Engine
arshadbarves
β€”
09-26 08:54amplified on r/LocalLLaMAreddit.post.1wqloas
jacek2023
peak 42 Β· 10 comments Β· 19% of case engagement
09-26 21:40amplified on hacker newshn.story.49860793
handfuloflight
peak 4 Β· 0 comments Β· 3% of case engagement
09-27 08:46amplified on r/MachineLearningreddit.post.1wrf4sm
bykof
peak 1 Β· 5 comments Β· 2% of case engagement
09-27 09:50amplified on hacker newshn.story.49865111
bykof
peak 2 Β· 0 comments Β· 1% of case engagement
09-27 09:59amplified on r/MachineLearningreddit.post.1wrgbhi
Spico197
peak 1 Β· 1 comments Β· 1% of case engagement
09-29 03:13amplified on hacker newshn.story.49887652
thebeardisred
peak 1 Β· 0 comments Β· 1% of case engagement
9 more amplifiers in ainews.case_chain
09-26 09:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wqloasβ€”
09-29 22:29reached heat=high Β· lag ? Β· via ledgerβ€”β€”
pace: p77 vs 1032 stories at the 336h mark (now 367h old) β€” ahead of cloudflare-security-audit-skill (1.0x), behind aisle-six-curl-cves (1.0x)

Evidence (16) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditinternlm/Intern-Decision 4B and 0.8B
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-26T09:23:39.538198+00:00

# [internlm](https://huggingface.co/internlm) / [Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) Like 1 Follow Intern Large Models 1.3k

[Image-Text-to-Text](https://huggingface.co/models?pipeline_tag=image-text-to-text)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)[qwen3\_5](https://huggingface.co/models?other=qwen3_5)[decision-making](https://huggingface.co/models?other=decision-making)[multimodal](https://huggingface.co/models?other=multimodal)[structured-prediction](https://huggingface.co/models?other=structured-prediction)[conversational](https://huggingface.co/models?other=conversational)

License: apache-2.0

[Model card](https://huggingface.co/internlm/Intern-Decision-4B)  [Files Files and versions  

xet](https://huggingface.co/internlm/Intern-Decision-4B/tree/main)  [Community](https://huggingface.co/internlm/Intern-Decision-4B/discussions)

 

Deploy

  Copy to bucket new   

Use this model

 

# Intern-Decision-4B

[Demo](https://huggingface.co/spaces/internlm/intern-decision) | [Model Weights](https://huggingface.co/collections/internlm/intern-decision) | [GitHub](https://github.com/internlm/Intern-Decision)

**Intern-Decision-4B** is a multimodal structured decision model fine-tuned from
**[Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)**.
It accepts a shared state, a schema of named questions, and optional images,
and returns an answer distribution for every question in one model forward pass.

## How inference works

1. Preserve the question and option order, and map each question's options to
   single-token symbols `A`, `B`, …, `Z`, `a`, …, `z`, `0`, …, `9`.
2. Render the original system prompt, state, decision schema, and a complete
   assistant JSON skeleton with one `<decision>` placeholder per field. Preserve
   the checkpoint's chat template and empty thinking block.
3. Run one causal Hugging Face forward pass. For the masked-next-token decision
   objective, read logits at the position **immediately before each placeholder**.
4. Take a softmax over only that field's allowed candidate-symbol logits, then
   apply the checkpoint's probability calibration.
5. Map symbols back to the original option values and return typed JSON answers.

This API performs structured candidate scoring. It does not call `generate()` or
sample free-form text. A request can contain multiple fields; no gold answers are
inserted into the prompt. The inference compiler uses only `state`, `questions`,
and optional `images`.

## Benchmark results

| Model | Jevbench-Easy | Jevbench-Original | Jevbench-Hard | Typed Decision | ToolACE | AG News | WildJailBreak | Average | Brier ↓ | ECE ↓ |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Jev | 100.00 | 98.61 | 72.07 | 73.35 | 91.29 | 89.57 | 96.29 | 88.74 | 0.358 | 0.095 |
| Laya | 95.83 | 72.22 | 28.83 | 35.95 | 63.87 | 92.84 | 14.84 | 57.77 | 0.804 | 0.246 |
| SemIf | 100.00 | 98.61 | 61.26 | 62.80 | 85.16 | 89.22 | 92.53 | 84.23 | 0.498 | 0.112 |
| Kev | 100.00 | 93.06 | 45.05 | 65.60 | 87.42 | 89.82 | 75.97 | 79.56 | 0.738 | 0.262 |
| JevK5 | 100.00 | 97.22 | 73.87 | 64.50 | 80.97 | 89.13 | 90.45 | 85.16 | 0.366 | 0.047 |
| Intern-Decision-0.8B | 97.92 | 80.56 | 52.25 | 77.35 | 94.52 | 88.61 | 64.48 | 79.38 | 0.530 | 0.066 |
| Intern-Decision-2B | 100.00 | 84.72 | 63.96 | 79.35 | 96.45 | 89.96 | 78.33 | 84.68 | 0.437 | 0.100 |
| Intern-Decision-4B | 100.00 | 98.61 | 73.87 | 80.55 | 96.45 | 90.82 | 89.86 | 90.02 | 0.347 | 0.065 |

## Inference latency

Measured on a single RTX 4090 with the local HF inference path. Values are
per-query end-to-end latency; they are workload and hardware dependent.

| Model | Mean | Median / P50 | P95 |
| --- | --- | --- | --- |
| Jev | 109.70 ms | 106.30 ms | 146.70 ms |
| Intern-Decision-0.8B | 33.98 ms | 33.44 ms | 37.50 ms |
| Intern-Decision-2B | 33.28 ms | 33.15 ms | 33.55 ms |
| Intern-Decision-4B | 44.16 ms | 44.03 ms | 44.60 ms |

## Known-distribution calibration pilot

This separate 96-case diagnostic uses exact reference distributions rather than
sampled hard labels. Lower is better. The pilot was not used to fit or select
the published temperature; the 4B model used its separately fitted T=1.992418.

| Category | Intern-Decision-4B before | Intern-Decision-4B after | Jev |
| --- | --- | --- | --- |
| Direct randomness and support | 0.483 / 0.181 | 0.421 / 0.129 | 0.490 / 0.216 |
| Composed events and mixtures | 0.677 / 0.254 | 0.577 / 0.150 | 0.682 / 0.274 |
| History, conditioning, and hidden state | 0.711 / 0.219 | 0.613 / 0.108 | 0.657 / 0.113 |
| Daily evidence and observation bias | 0.701 / 0.328 | 0.575 / 0.210 | 0.603 / 0.114 |
| Selective disclosure and probability puzzles | 0.540 / 0.119 | 0.510 / 0.049 | 0.483 / 0.138 |
| Sequential and combinatorial processes | 0.656 / 0.180 | 0.605 / 0.058 | 0.657 / 0.116 |
| **Overall (Brier / ECE)** | **0.628 / 0.213** | **0.550 / 0.089** | **0.595 / 0.130** |

## Quick start

Use **Python 3.12+**. Install `requirements.txt` in a suitable PyTorch/CUDA
environment, then import `DecisionEngine` from the downloaded model directory:

```
pip install -r requirements.txt
```

```
from inference import DecisionEngine

engine = DecisionEngine(device="cuda")  # Load once; reuse for subsequent requests.
request = {
    "state": "The customer was charged twice and asks for the extra payment back.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "billing": "Payments and refunds",
                "delivery": "Shipping and delivery",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "Rate the priority.",
            "criteria": ["Low", "Medium", "High"],
        },
        "refund_requested": {
            "type": "noul",
            "instructions": "Is the customer asking for a refund?",
        },
    },
}
response = engine.predict(request)  # One Python dict in, one response dict out.
print(response["answers"])
```

`predict(request)` accepts one request dictionary per call and returns a
JSON-serializable Jev-compatible response. It does not read request files or mutate
the supplied dictionary. Reuse the engine for each subsequent request.

The engine defaults to the checkpoint next to `inference.py`. To load another
local copy of this same model, use `DecisionEngine(checkpoint="./model-copy")`.
Use the inference module shipped with the selected size so its default calibration
matches. `backend="hf"` is the default and the only implemented backend. The
optional request `model` field does not switch checkpoints; the response `model`
identifies the weights actually loaded by this module.

### Request format

```
{
  "state": "The customer was charged twice and asks for the extra payment back.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this request?",
      "criteria": {
        "billing": "Payments and refunds",
        "delivery": "Shipping and delivery"
      }
    },
    "urgency": {
      "type": "score",
      "instructions": "Rate the priority.",
      "criteria": ["Low", "Medium", "High"]
    },
    "refund_requested": {
      "type": "noul",
      "instructions": "Is the customer asking for a refund?"
    }
  }
}
```

- **choice**: `criteria` is an ordered object mapping option values to descriptions.
- **score**: `criteria` is a list (values become `"0"`, `"1"`, …) or an ordered
  object with finite numeric string keys.
- **noul**: a binary decision with options `no`, then `yes`. Optional criteria can
  describe these values using `no`/`yes` or `false`/`true` keys.

Supply 1–16 questions, with up to 62 options per question. Inputs exceeding
`DecisionEngine(max_length=8192)` (default 8192 tokens) are rejected without truncation.

### Images

Set the request dictionary's `images` list in the intended order:

```
request["images"] = ["images/frame-1.png", "images/frame-2.png"]
response = engine.predict(request)
```

The checkpoint processor handles image resizing and token expansion. Relative
paths are resolved against `DecisionEngine(media_root=".")` (default: the working directory).
Supply up to eight images; image tokens count toward the input length limit.

### Response format

`answers` maps each field name to:

| Field | Meaning |
| --- | --- |
| `type` | `choice`, `score`, or `noul` |
| `probabilities` | Calibrated distribution over the original option values |
| `confidence` | Maximum candidate probability |
| `decision` | Highest-probability option value; lexical tie-breaking |
| `choice` | Selected value, for choice questions |
| `noul` | Probability of `yes`, for binary questions |
| `score` | Probability-weighted expected numeric value, for score questions |
| `legend` | Score values and their descriptions, for score questions |
| `source` | `local` |

The response follows the Jev envelope: `model`, `answers`, and `usage`.
It also includes `backend`, `timing`, and `calibration` as extension fields.
`usage.output_tokens` and `usage.decision_count` count scored fields, not generated
text tokens. `confidence` for a score question belongs to its most likely category;
the reported expected `score` can lie between categories.

## Calibration

The default temperature is **1.99241824**. It was fitted separately for this checkpoint
by NLL minimization on 1,728 designated calibration cases, with 1,693 separate
validation cases. Test-suite labels were not used to select the temperature.

The script follows the demo's numerical sequence:

```
p = softmax(candidate_logits.float())
calibrated_p = softmax(log(p) / T)
```

This is candidate probability calibration, **not a sampling temperature**. It
updates confidence, the `noul` probability, and the expected `score` while
preserving the argmax decision. For uncalibrated candidate probabilities, use
`DecisionEngine(temperature=1)`. A custom temperature must be finite and positive.

## License and acknowledgment

Intern-Decision is derived from the Qwen3.5 series. The original Qwen license is preserved
as [LICENSE-QWEN](https://huggingface.co/internlm/Intern-Decision-4B/tree/main/LICENSE-QWEN). Retain the license and applicable
upstream notices when redistributing. These weights were modified by decision
tuning, and this release adds the structured inference wrapper and model card.
We thank the Qwen team for the original models and multimodal processor.

Downloads last month
:   -

 

Safetensors

Model size

5B params

Tensor type

F32

Β·

BF16

Β·

Chat template

Files info

  

## Model tree for internlm/Intern-Decision-4B

Base model

[Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base)

Finetuned

  [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)

Finetuned

 ([784](https://huggingface.co/models?other=base_model:finetune:Qwen/Qwen3.5-4B)) 

this model
jacek20234110
🟧 echo.github ⭐Model card: Intern-Decision 4B/0.8B are multimodal structured decision models fine-tuned from Qwen3.5 that 'return an answer distribution foInternLM (internlm)β€”β€”
🟧 hnJulia-1: decision model that runs on almost anythinghandfuloflight40
🟧 hnShow HN: Peekaboolean – image and jev-like typed questions in typed anwers outbykof20
🟠 redditI trained a 500M VLM that answers typed questions about an image (choice / score / yes-no) with calibrated probabilities. ~400 ms on an M1 Pro, no text generation [P]
MachineLearning
bykof15
🟠 redditBetter, Faster, and More Calibrated than Jev, with Multi-Modal Ability [P]
MachineLearning
Spico19711
🟧 hnRun Decision Models on vLLM and Red Hat AI Using DiffusionGemmathebeardisred10
🟧 hnd1: Liquid AI's First Decision Modelmfiguiere30
🟧 hnLiquid AI releases decision model D1mnewme10
🟠 redditstuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model
LocalLLaMA
Inevitable-Log541410
🟠 redditJeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38Γ— faster decisions and +8.7 points accuracy, for under 2 GB extra memory
LocalLLaMA
Usual_Maximum767311432
🟧 hnFree hosted API for Laya, the open-weight decision modelboundlesshq11
🟠 redditbuilding an open source coding agent called Z-Engine
OpenAI
arshadbarves02
🟧 hnNew in Llama.cpp: Decision Modelscuvinny61
🟠 redditDecisionTune 1.0: a 395M encoder that picks from your options offline, about 10 ms per short decision on MLX (Apache-2.0)
LocalLLaMA
Abe238141
🟠 redditARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti
LocalLLaMA
KMatysek22

Interpretation history

Decision trace