2026-10-11 17:12 UTC

Independent reproduction will determine whether Qwen3.8-27B can run at 256K context on a 24GB RTX PRO 4000 SFF while achieving roughly 50 output tokens per second with MTP.

state: expiredheat: lowuncertainty: mediumconvergesscott: mediumqwen38 local-inference inference-economicsQwenNVIDIA

What is this?

The case concerns a reported local-inference result for Qwen3.8-27B on a 24GB NVIDIA RTX PRO 4000 SFF: a full 256K-token context and roughly 50 output tokens per second using multi-token prediction (MTP). The supplied snippets support that a four-bit 27B model may fit in 24GB only with very limited memory headroom, but one source explicitly says 24GB is plausible only at moderate context—not the full 262K window—and that no independent reproduction was available by its cutoff. Thus the exact 256K/50 tok/s result remains an unverified performance claim, with quantization, KV-cache placement, workload, and measurement conditions not established here.

Why it matters to Scott

If independently reproduced, the result would extend Scott’s hardware-aware local-inference work by showing that a quantized 27B model can combine a 256K resident context with interactive MTP throughput on a constrained 24GB workstation GPU—potentially actionable for gamepc. It currently remains seller/testimony-grade evidence rather than a validated capability, matching Scott’s Capability Audit and Evidence Class Ladder requirements; the radar tracks closely related long-context, KV-cache, and MTP memory questions but not this exact result.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.capability-auditip:concept.evidence-class-ladderradar:concept.local-inferenceradar:concept.long-context-inferenceradar:concept.kv-cacheradar:concept.speculative-decodingradar:llama-cpp-mtp-autofit-memory
queries asked of Scott's wikis
  • 24GB local inference and long-context economics
  • KV-cache memory limits at 256K context
  • multi-token prediction decode throughput
  • quantization tradeoffs for local 27B models
  • independent benchmarking of local inference claims
  • workstation GPUs versus cloud inference economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (33) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnQwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTPpich3634
🟧 echo.blog ⭐Reports running Qwen3.8-27B at 256K context on a 24GB RTX PRO 4000 SFF and obtaining approximately 50 tokens per second with MTP.pich——
🟠 redditMy experience with Qwen 3.8 27B Q3_XL on an RTX 3090
LocalLLaMA
cezarducatti713
🟠 redditArtificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
LocalLLaMA
anderspitman1137433
🟠 redditQwen3.8-27B lands next to DeepSeek V4 and GPT-5.6 Luna Max on the Artificial Analysis Benchmark. You can now run a near frontier model with just a RTX 3090.
singularity
yaboyyoungairvent544113
🟧 hnQwen3.8 27B scores 52 on Artificial Analysisanana_370174
🟠 redditQwen3.8 (27b) outperforms GPT-5.6-Terra (Max) for Agentic tasks!
LocalLLaMA
UnknownEssence69
🟠 redditQwen3.8-27B for the RAM Poor Mac user:
LocalLLaMA
JLeonsarmiento30
🟠 redditHow are you hosting Qwen3.8-27b with a 5090?
LocalLLaMA
DustNearby2848126
🟠 redditLLMs Endgame: This is Unreasonable
LocalLLaMA
Potential_Block45982238
🟠 redditI pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090
LocalLLaMA
iamMess3116
🟠 redditQwen 3.8 27B - RTX 4090 24Gb - Sharing my Config
LocalLLaMA
gavwhittaker132
🟠 redditThis is what context management still is, Qwen 3.8 27B?
LocalLLaMA
juss-i05
🟧 hnI pushed Qwen3.8-27B to 99 tps on a single request and 1150 on batch on a 3090mess133711
🟠 redditBenchmarked Qwen3.8-27B on 4x RTX 3090
LocalLLaMA
Mr_Moonsilver1418
🟠 redditQwen3.8-27B (Q6_K_XL) speeds on M2 Ultra 192GB — what are you getting on your Ultra?
LocalLLaMA
planetearth80218
🟠 redditwe benchmark models nobody actually runs
LocalLLaMA
AuspiciousApple7070
🟧 hnQwen3.8-27B on a single RTX 3090: crash fix, 131K context, 9 mythsjonaddb10
🟠 redditQwen 3.8 q4kxl made by UNSLOTH is collapsing more or less over a 60k tokens
LocalLLaMA
Healthy-Nebula-3603713
🟠 redditQwen3.8-27B at 5.01 BPW: 256K context, Q4_1-level PPL and 50.44 tok/s on a 24 GB Blackwell
LocalLLaMA
iam313371211
🟧 hnWe Tested Qwen3.8 27B: How Much GPU and VRAM Do You Needmdp202110
🟠 redditQwen3.8 27b: Speed As Context Grows - 1,274 generations, RTX 5090
LocalLLaMA
_-_David214
🟠 reddit[Guide] Squeezing ~18–20 tok/s out of Qwen3.8-27B on 16GB VRAM + 64GB System RAM (Without sacrificing KV Cache quality!)
LocalLLaMA
BassAzayda104
🟠 reddit[Guide] Squeezing Qwen3.8-27B (256k Context) onto a Single 16GB GPU (4070 Ti Super) — 100% VRAM Offload + N-Gram Speculative Decoding
LocalLLaMA
ndiphilone516
🟠 redditThinking of opening free Qwen3.8-27B access to the community for few days
LocalLLaMA
No_Run8812027
🟠 redditOptimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
LocalLLaMA
MaxDev03117
🟠 redditQwen 3.8 27B is faster than expected
LocalLLaMA
chocofoxy3055
🟠 redditQwen3.8 Q4_k_m 1M context Strix Halo / 3080 -> 45 tps
LocalLLaMA
TrifleHopeful541846
🟠 redditAnyone running qwen 3.8 27b on 5070ti (16GB)?
LocalLLaMA
zannix1759
🟠 redditQwen3.8-27B: slower tokens, faster and better results
LocalLLaMA
surreal_tournament302116
🟠 redditI pushed Qwen3.8-27B to 124 tps on a single request on a RTX 3090
LocalLLaMA
iamMess13774
🟠 redditQwen3.8-27B VRAM on 16GB with 50tok/s, 85k q8 context
LocalLLaMA
brainExploded991334
🟠 redditPessimistic electricity cost calculations
LocalLLaMA
Thin_Pollution8843638

Interpretation history

Decision trace