2026-10-11 16:37 UTC

Inception claims Mercury 2.5 Preview uses diffusion-style generation to provide sufficiently low-latency language-model inference for interactive and agent workloads, potentially offering an alternative to conventional autoregressive serving.

state: acceleratingheat: mediumuncertainty: mediumconvergesscott: mediumfrontier-models inference-economicsInceptionOpenRouter

What is this?

Mercury 2.5 Preview is a diffusion language model from Inception Labs that iteratively refines tokens in parallel rather than decoding strictly left-to-right; it is offered through Inception’s API, OpenRouter, and Baseten for latency-sensitive uses such as coding, voice, search, and agents. Inception advertises a 260K context window, reasoning, parallel tool calls, structured JSON, and 1,107 tokens/sec on commonly available NVIDIA GPUs, while the case reports an independent Artificial Analysis result near 770 tokens/sec that corroborates unusually high raw throughput. The supplied material does not establish matched-quality end-to-end latency, dependable tool execution, concurrency behavior, or lower cost per successful agent task versus strong autoregressive models.

Why it matters to Scott

Mercury’s independently measured raw throughput converges with Scott’s position that latency is binding in interactive agent loops and makes the model an actionable candidate for his OpenRouter-based routing and trace-backed evaluation work. It could alter model selection, but matched-quality agent reliability and cost per successful task still need workload-specific testing before it bears on production architecture.
ip:concept.latencyip:concept.real-time-ai-systemsip:concept.evaluation-driven-developmentip:concept.ai-unit-economicsdev:project.remote-execdev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisondev:technology.openrouterradar:concept.diffusion-language-modelsradar:concept.inference-latencyradar:concept.inference-economicsradar:diffusiongemma-language-model-validation
queries asked of Scott's wikis
  • agent-loop latency as a binding constraint
  • diffusion language models and parallel decoding
  • model routing across quality speed and cost
  • cost per successful agent task
  • tool-use reliability and structured-output evaluation
  • OpenRouter model evaluation harnesses

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1010h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-30 14:00⭐ origin echo-reconstructedThe earliest accessible first-party artifact is Inception Labs’ official model page, which presents “Mercury 2.5 (Preview)” as its “most int
Inception Labs on other (echo) · attributed from hn.story.49537813
—
09-02 15:31first on hacker news · published · +73.5hUltra Fast LLM's Are Interesting (Mercury 2.5 Preview Release)
gmarkwa
—
09-25 03:23first on r/artificial · published · +613.4hStefano Ermon: Autoregressive inference is sequential and memory-bound. Diffusion is built to map to GPUs — that's why it wins.
cen6wkf
—
09-02 15:31amplified on hacker newshn.story.49537813
gmarkwa
peak 1 · 1 comments · 0% of case engagement
09-04 06:16amplified on hacker newshn.story.49561124
E-Reverance
peak 3 · 0 comments · 1% of case engagement
09-08 16:44amplified on hacker newshn.story.49612827
fittingopposite
peak 1 · 1 comments · 0% of case engagement
09-08 20:14amplified on hacker news 👑hn.story.49616354
Topfi
peak 248 · 54 comments · 54% of case engagement
09-19 10:16amplified on hacker newshn.story.49765134
peter_d_sherman
peak 2 · 1 comments · 1% of case engagement
09-23 22:16amplified on hacker newshn.story.49823348
Retro_Dev
peak 151 · 92 comments · 44% of case engagement
1 more amplifiers in ainews.case_chain
09-02 16:21our radar first saw it · +74.4hdiscovery anchor: hn.story.49537813—
pace: p84 vs 519 stories at the 720h mark (now 1010h old) — ahead of ling-spark-mtp-throughput (1.0x), behind compute-cheap-h100-h200-pricing (1.0x)

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnUltra Fast LLM's Are Interesting (Mercury 2.5 Preview Release)gmarkwa11
🟧 echo.other ⭐The earliest accessible first-party artifact is Inception Labs’ official model page, which presents “Mercury 2.5 (Preview)” as its “most intInception Labs——
🟧 hnUnlocking Lossless Speedups in LLMs via Discrete Diffusion (5000 Tk/S)E-Reverance30
🟧 hnInception Launches Mercury 2.5, the Next Tier of Intelligence for Diffusion LLMsfittingopposite11
🟧 hnMercury 2.5Topfi24854
🟧 hnBuilding the fastest LLMs: why we're starting with diffusionpeter_d_sherman21
🟧 hnMercury 2.5 LLM hits 770 tokens per secondRetro_Dev15192
🟠 redditStefano Ermon: Autoregressive inference is sequential and memory-bound. Diffusion is built to map to GPUs — that's why it wins.
artificial
cen6wkf03

Interpretation history

Decision trace