2026-10-11 17:10 UTC

Independent benchmarks will determine whether Ninfer delivers competitive throughput, reliability, and memory efficiency for its supported model checkpoints and single-GPU configurations.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference inference-economics ai-infrastructureNeroued
Surfaced 2026-09-21T15:44:11Z — A high-performance single-GPU inference implementation for selected model checkpoints and GPUs. — Ninfer has moved from isolated speed claims into an expanding implementation ecosystem: hardware ports, production-like agent use, and now a native typed-decision API proof of concept. The loud cross-platform spread and continuing derivatives warrant immediate attention, although no quality-matched benchmark yet establishes a general advantage over vLLM or llama.cpp.

What is this?

Ninfer is a deliberately narrow, from-scratch C++/CUDA single-GPU inference engine by developer Neroued (Apache-2.0, first committed June 26, 2026) that trades generality for throughput: it serves Qwen3.5 dense/MoE checkpoints on a single RTX 5090 with a startup-fixed 1–8 concurrent requests via CLI and OpenAI-/Anthropic-compatible HTTP APIs, claiming ~700 tok/s decode and ~15.5k tok/s prefill, and unusually ships a 225-response quality audit with published benchmark tables (e.g., Qwen3.8-27B NVFP4: 96.67% AIME 2025/2026, 90.4% GPQA-Diamond, 83.53% RealWorldQA) and MTP3/DFlash2 concurrency-scaling results with reproduction commands. The web snippets are almost entirely first-party — the project's own repo/docs and an NYU Shanghai library writeup — plus an early Reddit comparison putting Ninfer at ~200 tok/s vs ~140 tok/s for llama.cpp q4km on a single request, with the author attributing speed to sm120-specific custom CUDA kernels while conceding output is not exactly 1:1 with llama.cpp quality. The contested competitive picture this case actually tracks — tuned vLLM builds winning agentic/concurrent workloads, exl3 leading compression-per-bit, heavy fork fragmentation across 4090/3090/Turing/dual-GPU ports, reliability caveats (cache-related TTFT degradation, tool-call leakage, Windows focus throttling), and MegaCapybara's unverified ~2x-decode claim — rests entirely on the case's own Reddit evidence trail, not these snippets; no independent, quality-matched end-to-end benchmark appears anywhere in the supplied web material.

Why it matters to Scott

Two months of community head-to-heads have independently arrived where Scott already argued: caching behavior, not headline tok/s, governs real agent-serving economics (Ninfer cache bugs producing seconds-to-minutes TTFT degradation, warm-cache prefill contamination flagged by critics), and allocated context is not useful context (1M-token forks with accuracy unvalidated) — dated field receipts for his prefix-caching-economics and attention-budget/context-rot positions rather than a challenge to them. The case also keeps his gamepc/Ollama/LiteLLM local serving decision warm through a contested regime map (Ninfer wins short-context bursts, tuned vLLM wins agentic/concurrent, exl3 wins compression-per-bit) and MegaCapybara's unverified 2x claim — exactly the quality-matched comparison his trace-backed-agent-comparison method exists to run once MegaCapybara's source lands, but nothing here is validated enough yet to change what he actually deploys.
dev:project.gamepcdev:technology.ollamadev:technology.litellmdev:concept.hardware-aware-local-inferencedev:concept.trace-backed-agent-comparisonip:concept.prefix-caching-economicsip:concept.attention-budgetip:concept.context-rotradar:ninfer-qwen-5090-throughputradar:concept.local-inferenceradar:concept.inference-benchmarkingradar:concept.speculative-decodingradar:concept.quantizationradar:concept.vllmradar:magnitude-self-optimizing-inference-engineradar:llamacpp-fork-fragmentationradar:qwen38-dflash2-long-context-speedupradar:qwen38-27b-24gb-long-context-throughput
queries asked of Scott's wikis
  • local inference engine choice for coding agents
  • single-GPU serving stack vLLM llama.cpp tradeoffs
  • quantization quality impact on agent task success
  • speculative decoding MTP acceptance agent workloads
  • long-context degradation useful context window
  • agent-serving reliability TTFT cache bugs benchmark methodology

Measured heat

now 0 pts/hpeak 17 pts/hcomments 0/hpeers p25momentum: steady3 platformsage 1376h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-15 08:27 (minted)⭐ origin echo-reconstructedA high-performance single-GPU inference implementation for selected model checkpoints and GPUs.
Neroued on github (echo) · attributed from hn.story.49308615 · published time unknown
—
08-15 07:45first on hacker news · published · lag ?High-performance single-GPU inference for selected model checkpoints and GPUs
davedx
—
08-15 19:12first on r/LocalLLaMA · published · lag ?Ninfer for RTX 4090 and Qwen 3.8 27B
UDPSendToFailed
—
08-15 07:45amplified on hacker newshn.story.49308615
davedx
peak 1 · 0 comments · 0% of case engagement
08-15 19:12amplified on r/LocalLLaMAreddit.post.1vpbdq3
UDPSendToFailed
peak 3 · 4 comments · 0% of case engagement
08-15 21:04amplified on r/LocalLLaMAreddit.post.1vpe2uw
Ond7
peak 45 · 27 comments · 4% of case engagement
08-16 19:38amplified on r/LocalLLaMAreddit.post.1vq6fdj
iamMess
peak 156 · 68 comments · 13% of case engagement
08-16 20:48amplified on r/LocalLLaMAreddit.post.1vq881r
UDPSendToFailed
peak 33 · 29 comments · 3% of case engagement
08-17 03:05amplified on r/LocalLLaMAreddit.post.1vqglk6
C1oover
peak 9 · 12 comments · 1% of case engagement
24 more amplifiers in ainews.case_chain
08-15 08:21our radar first saw it · lag ?discovery anchor: hn.story.49308615—
09-21 15:44reached heat=high · lag ? · via ledger——

Evidence (31) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnHigh-performance single-GPU inference for selected model checkpoints and GPUsdavedx10
🟧 echo.github ⭐A high-performance single-GPU inference implementation for selected model checkpoints and GPUs.Neroued——
🟠 redditNinfer for RTX 4090 and Qwen 3.8 27B
LocalLLaMA
UDPSendToFailed34
🟠 reddit880 tok/s on one 5090 Qwen3.8-27B in 4-bit NVFP4, full 262k context
LocalLLaMA
Ond74127
🟠 redditQwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak
LocalLLaMA
iamMess15668
🟠 redditNInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM
LocalLLaMA
UDPSendToFailed3329
🟠 redditWhat happened to exl3/tabbyapi?
LocalLLaMA
C1oover912
🟠 redditQuestion about making local LLMs faster
LocalLLaMA
Viktri138
🟠 reddit45 tok/s Qwen3.8-27B MTP3 on modded RTX 2080 Ti 22GB with my NInfer port
LocalLLaMA
xrailgun517
🟠 redditOrnith1.5/Qwen3.8 Ninfer & omlx quants
LocalLLaMA
Pyros-SD-Models57
🟠 redditSharp template to NInfer: -42% output tokens, same speed
LocalLLaMA
xrailgun2921
🟠 redditI forked Ninfer 3090 and converted it to run on the CMP170HX - doubled my Qwen3.6-35B from llama.cpp
LocalLLaMA
ubrtnk2422
🟠 redditNInfer 4090 Windows update is out with 1.5-2k t/s prefill, extended MTP, disk caching with DirectStorage, built-in llama.cpp WebUI and more
LocalLLaMA
UDPSendToFailed1831
🟠 redditAre the best settings for single 3090 just ninfer-3090 build or can i do better?
LocalLLaMA
randomjapaneselearn421
🟠 redditNinfer on a 5090 & Qwen3.8 27B w/ Vision any tips for it to have decent context?
LocalLLaMA
Rollingsound514224
🟠 redditNinfer and a 5090 with 3.8 27B is making me cry tears of joy it's so good.
LocalLLaMA
Rollingsound514187161
🟠 redditQwen3.8-27B on a 24GB RTX PRO 4000 Blackwell: 128K real context, 785 tok/s prefill, 67 tok/s MTP3 decode
LocalLLaMA
mmkaywhatevers1518
🟠 reddit(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s
LocalLLaMA
Littlepharaoh2741
🟠 redditQwen 3.8 27B NVFP4, in a single 5090, using nInfer above 200tps at 180K contexts
LocalLLaMA
Maleficent-Ad5999015
🟠 redditWhich is better ninfer vs vllm for Qwen 3.8 27B on RTX 5090?
LocalLLaMA
MaxKingCS627
🟠 redditQwen3.8-27B on 2× RTX 5070 Ti 16GB — llama.cpp vs vLLM vs NInfer benchmarks
LocalLLaMA
puthre1116
🟠 redditNInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090
LocalLLaMA
bengizmoed8865
🟠 redditMy local LLM demoscene generator can now watch its own output and rewrite it!
LocalLLaMA
jacobpederson71
🟠 redditNInfer fork: 555k context@fp4 for 5090 with YARN, reliable kv host cacheing, monitoring, jinja, opened model support
LocalLLaMA
Lumpy-Comedian-10272160
🟠 redditFor the brave: Ninfer + MTP + Vision + 400k context (nvfp4 quant and kv)
LocalLLaMA
DavidOelfke27
🟠 redditSolved: LLM inference on Windows was 2–3x slower when the server window wasn't focused
LocalLLaMA
koloved1325
🟠 redditninfer-3090 single thread mini-benchmark results
LocalLLaMA
milkipedia911
🟠 reddit600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?
LocalLLaMA
CharlesStross10468
🟠 redditGot jev-like api running natively on ninfer / qwen3.8 27b and results are quite decent
LocalLLaMA
Unlucky-Message8866113
🟠 redditThe fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.
LocalLLaMA
BringTea_666063
🟠 redditRunning the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s
LocalLLaMA
Distinct-Pie23891333

Interpretation history

Decision trace