A LocalLLaMA builder claims a custom CUDA megakernel fusing entire speculative-decoding cycles achieves 1.4โ1.9ร speedups over llama.cpp for Qwen3.8-27B on a single RTX 3090, with the kernel written using Claude Opus 5.5.
state: watchingheat: highuncertainty: mediumnovelscott: highlocal-inference speculative-decoding custom-kernels cuda qwen3.8-27bAdorable_Weakness_39
What is this?
A Reddit user 'Adorable_Weakness_39' on r/LocalLLaMA claims to have written a custom CUDA megakernel (using Claude Opus 5.5) that fuses entire speculative-decoding cycles for Qwen3.8-27B, achieving 1.4โ1.9ร speedups over llama.cpp on a single RTX 3090 (~140 tok/s on code). The web results confirm active community benchmarking of speculative decoding on Qwen 3.6-series models on 3090s (with llama.cpp's built-in n-gram/MTP spec-decoding showing ~75% acceptance and up to 2ร throughput on dense 27B), but do not mention this specific builder, the Qwen3.8-27B variant, a custom megakernel, or Claude Opus 5.5 authorship. One controlled benchmark (thc1006) found no net speedup for the 35B-A3B MoE variant on Ampere. The claim is concrete and measurable but currently uncorroborated by the supplied snippets.
Why it matters to Scott
This builder claim directly bears on Scott's hardware-aware local inference concept (dev:concept.hardware-aware-local-inference) โ a custom megakernel fusing entire speculative-decoding cycles is the archetype of treating accelerator placement, memory pressure, and compilation as explicit runtime policy. It also converges on his custom-kernel strategy (ip:framework.worldview-recursive-compression, ip:concept.kernel, ip:concept.kernel-flywheel), AI-assisted CUDA development workflows (dev:technology.claude-code, ip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebook, ip:framework.code-first-architecture), model sovereignty / local-first stack preferences (ip:framework.sovereign-software-assurance, dev:project.gamepc, dev:technology.ollama), and single-GPU inference economics for 27B-class models (ip:framework.the-mature-token-law, ip:concept.ai-unit-economics). The radar is already tracking adaptive speculative decoding on a $300 consumer GPU (radar:adaptive-speculative-decoding-300-gpu) and another LocalLLaMA builder's controlled pipeline comparison (radar:agentic-retrieval-frames-benchmark), making this a concrete instance of a pattern the radar watches.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:framework.worldview-recursive-compressionip:concept.kernelip:concept.kernel-flywheeldev:technology.cudadev:technology.claude-codeip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookip:source.agentic-coding-plain-and-spicy-ebookip:framework.code-first-architectureip:concept.code-as-step-between-model-runsip:framework.sovereign-software-assuranceip:concept.inference-time-scalingip:framework.the-mature-token-lawip:concept.ai-unit-economicsradar:adaptive-speculative-decoding-300-gpuradar:agentic-retrieval-frames-benchmarkradar:adaptive-kv-cache-streamingradar:amd-llama-cpp-prefill-speedupradar:ante-offline-coding-agentradar:agentic-animation-production-patternradar:afm3-prompt-conditioned-pruningradar:applied-compute-training-serving-platformradar:0pirate-ast-anonymizer-mcp-proxyradar:aetheris-code-first-cad-kernel
queries asked of Scott's wikis
- local-inference custom-kernel strategy and builder tooling
- speculative-decoding implementation patterns in llama.cpp vs custom engines
- AI-assisted CUDA kernel development (Claude/Opus) as a workflow
- single-GPU inference economics for 27B-class models
- model sovereignty / local-first inference stack preferences
Measured heat
now 4 pts/hpeak 27 pts/hcomments 3/hpeers p66momentum: steady1 platformsage 27h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p80 vs 968 stories at the 24h mark (now 27h old) โ ahead of little-gemma-jetson-voice-inference (1.0x), behind rails-cve-hours-to-exploitation (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-11T14:21:45Z
Follow-up accuracy benchmarks (KL divergence 0.0009 vs llama.cpp, perplexity parity at 96k context) address the main fidelity concern for the megakernel. Community engagement is accelerating (79th-percentile velocity, expanding discussion with quantization-port requests and 3090 owners seeking reproduction). Scott's explicit up-vote confirms high relevance. Still single-source with no independent replication, but the claim is now auditable on both speed and quality dimensions.
2026-10-11T13:39:44Z
evidence attached: reddit.post.1x37q7b โ Follow-up accuracy benchmarks (KL divergence 0.0009 vs llama.cpp) for the custom CUDA megakernel claimed in the open case.
2026-10-11T04:40:40Z
Scott's explicit up-vote confirms high relevance to his hardware-aware local inference and custom-kernel frameworks. Velocity spikes (6โ10ร baseline) and 83rd-percentile peer engagement show the builder community is actively discussing the claim. The case remains a single-source claim with published code but no independent replication yet; the periphery is expanding (discussion, quantisation requests, 3090 owners asking for reproduction) but each addition is still thin.
2026-10-10T13:55:14Z
grounded: novel/high โ This builder claim directly bears on Scott's hardware-aware local inference concept (dev:concept.hardware-aware-local-inference) โ a custom megakernel fusing en
2026-10-10T13:41:51Z
case created โ Concrete builder claim with measured speedups on specific model/hardware; distinct from existing custom-engine cases.
Decision trace
- 10-12 01:29attention_routeNew accuracy benchmarks (KL divergence, perplexity) resolve the key open question about fidelity; Scott up-voted (feedback_interrupt) signaling high priority; scheduled for next briefing due to quiet
- 10-12 01:21attention_candidatematerial_reprice
- 10-12 01:21repriceFollow-up accuracy benchmarks (KL divergence 0.0009 vs llama.cpp, perplexity parity at 96k context) address the main fidelity concern for the megakernel. Community engagement is accelerating (79th-per
- 10-12 00:43attention_routeNew accuracy benchmarks (KL divergence, perplexity) address the main open question about the megakernel's fidelity; user feedback_interrupt signals high interest; scheduled for next briefing due
- 10-12 00:39attention_candidateattach
- 10-12 00:39attachFollow-up accuracy benchmarks (KL divergence 0.0009 vs llama.cpp) for the custom CUDA megakernel claimed in the open case.
- 10-12 00:36propose_attachFollow-up accuracy benchmarks (KL divergence 0.0009 vs llama.cpp) for the custom CUDA megakernel claimed in the open case.
- 10-11 15:40repriceScott's explicit up-vote confirms high relevance to his hardware-aware local inference and custom-kernel frameworks. Velocity spikes (6โ10ร baseline) and 83rd-percentile peer engagement show the
- 10-11 15:27sensor_dirtyvelocity_spike
- 10-11 15:00feedback_interruptScott vote via UI
- 10-11 14:27sensor_dirtycomment_update
- 10-11 10:03attention_communicatedA builder released a custom CUDA megakernel that fuses entire speculative-decoding cycles into a single kernel launch, achieving 1.4โ1.9ร speedups over llama.cpp with MTP for Qwen3.8-27B (unsloth Q4_K
- 10-11 10:03attention_routeHigh-relevance builder artifact with published code and measurable speedups converges on multiple load-bearing frameworks (hardware-aware local inference, kernel flywheel, sovereign software assurance
- 10-11 08:31sensor_dirtyvelocity_spike
- 10-11 05:37sensor_dirtycomment_update
- 10-11 01:21attention_routeHigh-relevance builder artifact with published code and measurable speedups converges on multiple load-bearing frameworks (hardware-aware local inference, kernel flywheel, sovereign software assurance
- 10-11 01:18attention_candidatecreate
- 10-11 00:55groundThis builder claim directly bears on Scott's hardware-aware local inference concept (dev:concept.hardware-aware-local-inference) โ a custom megakernel fusing entire speculative-decoding cycles is
- 10-11 00:41createConcrete builder claim with measured speedups on specific model/hardware; distinct from existing custom-engine cases.