2026-10-11 16:35 UTC

A LocalLLaMA builder claims a custom CUDA megakernel fusing entire speculative-decoding cycles achieves 1.4โ€“1.9ร— speedups over llama.cpp for Qwen3.8-27B on a single RTX 3090, with the kernel written using Claude Opus 5.5.

state: watchingheat: highuncertainty: mediumnovelscott: highlocal-inference speculative-decoding custom-kernels cuda qwen3.8-27bAdorable_Weakness_39

What is this?

A Reddit user 'Adorable_Weakness_39' on r/LocalLLaMA claims to have written a custom CUDA megakernel (using Claude Opus 5.5) that fuses entire speculative-decoding cycles for Qwen3.8-27B, achieving 1.4โ€“1.9ร— speedups over llama.cpp on a single RTX 3090 (~140 tok/s on code). The web results confirm active community benchmarking of speculative decoding on Qwen 3.6-series models on 3090s (with llama.cpp's built-in n-gram/MTP spec-decoding showing ~75% acceptance and up to 2ร— throughput on dense 27B), but do not mention this specific builder, the Qwen3.8-27B variant, a custom megakernel, or Claude Opus 5.5 authorship. One controlled benchmark (thc1006) found no net speedup for the 35B-A3B MoE variant on Ampere. The claim is concrete and measurable but currently uncorroborated by the supplied snippets.

Why it matters to Scott

This builder claim directly bears on Scott's hardware-aware local inference concept (dev:concept.hardware-aware-local-inference) โ€” a custom megakernel fusing entire speculative-decoding cycles is the archetype of treating accelerator placement, memory pressure, and compilation as explicit runtime policy. It also converges on his custom-kernel strategy (ip:framework.worldview-recursive-compression, ip:concept.kernel, ip:concept.kernel-flywheel), AI-assisted CUDA development workflows (dev:technology.claude-code, ip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebook, ip:framework.code-first-architecture), model sovereignty / local-first stack preferences (ip:framework.sovereign-software-assurance, dev:project.gamepc, dev:technology.ollama), and single-GPU inference economics for 27B-class models (ip:framework.the-mature-token-law, ip:concept.ai-unit-economics). The radar is already tracking adaptive speculative decoding on a $300 consumer GPU (radar:adaptive-speculative-decoding-300-gpu) and another LocalLLaMA builder's controlled pipeline comparison (radar:agentic-retrieval-frames-benchmark), making this a concrete instance of a pattern the radar watches.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:framework.worldview-recursive-compressionip:concept.kernelip:concept.kernel-flywheeldev:technology.cudadev:technology.claude-codeip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookip:source.agentic-coding-plain-and-spicy-ebookip:framework.code-first-architectureip:concept.code-as-step-between-model-runsip:framework.sovereign-software-assuranceip:concept.inference-time-scalingip:framework.the-mature-token-lawip:concept.ai-unit-economicsradar:adaptive-speculative-decoding-300-gpuradar:agentic-retrieval-frames-benchmarkradar:adaptive-kv-cache-streamingradar:amd-llama-cpp-prefill-speedupradar:ante-offline-coding-agentradar:agentic-animation-production-patternradar:afm3-prompt-conditioned-pruningradar:applied-compute-training-serving-platformradar:0pirate-ast-anonymizer-mcp-proxyradar:aetheris-code-first-cad-kernel
queries asked of Scott's wikis
  • local-inference custom-kernel strategy and builder tooling
  • speculative-decoding implementation patterns in llama.cpp vs custom engines
  • AI-assisted CUDA kernel development (Claude/Opus) as a workflow
  • single-GPU inference economics for 27B-class models
  • model sovereignty / local-first inference stack preferences

Measured heat

now 4 pts/hpeak 27 pts/hcomments 3/hpeers p66momentum: steady1 platformsage 27h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-10 13:01โญ origin directly observedQwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
Adorable_Weakness_39 on r/LocalLLaMA
โ€”
10-11 13:04first on r/LocalLLaMA ยท published ยท +24.1hUPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
Adorable_Weakness_39
โ€”
10-10 13:01amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1x2erdj
Adorable_Weakness_39
peak 94 ยท 54 comments ยท 86% of case engagement
10-11 13:04amplified on r/LocalLLaMAreddit.post.1x37q7b
Adorable_Weakness_39
peak 17 ยท 7 comments ยท 14% of case engagement
10-10 13:33our radar first saw it ยท +0.5hdiscovery anchor: reddit.post.1x2erdjโ€”
10-11 14:21reached heat=high ยท +25.3h ยท via queue+ledgerโ€”โ€”
pace: p80 vs 968 stories at the 24h mark (now 27h old) โ€” ahead of little-gemma-jetson-voice-inference (1.0x), behind rails-cve-hours-to-exploitation (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญQwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
LocalLLaMA
Adorable_Weakness_399454
๐ŸŸ  redditUPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
LocalLLaMA
Adorable_Weakness_39177

Interpretation history

Decision trace