Nori claims its released LLM serving stack sustains over one million tokens per second in reproducible workloads, which if validated would establish a new high-throughput inference serving point and materially shift serving economics.
state: seedheat: lowuncertainty: mediumnovelscott: mediumllm-serving inference-throughputNori (Noriagentic)
What is this?
Nori (Noriagentic) has released an LLM serving stack and, per its own post echoed through an HN submission titled 'Nori LLM: Achieving Over 1M tok / s', claims the stack sustains over one million tokens per second on reproducible workloads. The supplied web material contains no independent coverage or validation of Nori: nothing establishes the hardware, model, batch size, or whether the figure is per-GPU, per-node, or fleet-aggregate โ and the distinction matters, since a million tok/s is routine at fleet scale but extraordinary as a per-node claim. The surrounding results only frame why throughput claims matter: inference-economics literature puts GPT-4-class serving at ~$0.40 per million tokens in 2026, identifies decode as memory-bandwidth-bound, and cites prior per-request maxima around 400 tok/s (Artificial Analysis 2024, Llama 3 8B โ a per-request metric, not aggregate serving throughput). The claim therefore stands as unvalidated first-party testimony with near-zero third-party traction so far.
Why it matters to Scott
If validated at per-node (not fleet-aggregate) scale, a 1M tok/s serving point would move the supply-side economics beneath his tokens-as-fuel frameworks โ the Agent Token Manifesto's premise, his unit-economics work, and the LLM pricing guide he published โ while the claim's current state (first-party testimony, measurement unit undefined, no independent traction) sits at the low rungs of his discussed-is-not-deployed evidence ladder. Until independent benchmarks land, it extends rather than changes what he argues, so the per-GPU-vs-per-node ambiguity is the thing to watch, not a premise to update yet.
ip:source.the-agent-token-manifestoip:concept.ai-unit-economicsip:source.discussed-is-not-deployed-ebookdev:project.llmreportradar:concept.llm-servingradar:concept.inference-economicsradar:concept.benchmark-integrityradar:concept.reproducibilityradar:concept.vllmradar:cpp-vllm-serving-port-validation
queries asked of Scott's wikis
- inference serving stack throughput vLLM batching self-hosted
- cost per million tokens inference economics GPU utilization
- agent harness token consumption throughput latency requirements
- benchmark reproducibility vendor claims AI infrastructure marketing
- local inference hardware selection tokens per second per GPU
- serving stack release open-source deployment engineering
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 420h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p57 vs 1032 stories at the 336h mark (now 420h old) โ ahead of 37signals-agent-driven-default (1.0x), behind heif-heist-parser-disclosure (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-09-24T04:36:22Z
grounded: novel/medium โ If validated at per-node (not fleet-aggregate) scale, a 1M tok/s serving point would move the supply-side economics beneath his tokens-as-fuel frameworks โ the
2026-09-24T04:29:02Z
case created โ A consequential first-party throughput claim squarely on Scott's inference-economics radar, worth a seed despite near-zero HN traction so far.
Decision trace
- 09-26 05:59review_screenjev screen: no material development (noul=0.06)
- 09-24 14:36groundIf validated at per-node (not fleet-aggregate) scale, a 1M tok/s serving point would move the supply-side economics beneath his tokens-as-fuel frameworks โ the Agent Token Manifesto's premise, hi
- 09-24 14:29createA consequential first-party throughput claim squarely on Scott's inference-economics radar, worth a seed despite near-zero HN traction so far.