2026-10-11 16:36 UTC

llama-cpp

band: hotmomentum: stable score: 1.0
temperature history

Episodes (50)

Independent benchmarks will determine whether llama.cpp PR 25940 reproducibly improves ROCm prompt processing by roughly 15% and fixes the reported 28-fold Q2_K slowdown across AMD GPU configurations.
expiredknownscott: medium
Independent use will determine whether Pi 0.81's native llama.cpp router materially simplifies local-model setup and operation for agent workflows compared with extensions or manual configuration.
expiredconvergesscott: medium
Independent benchmarks will determine whether Poolside's 120B-class Laguna-S 2.1 is competitive for coding and practical local inference.
resolvedknownscott: medium
Independent benchmarks and mainstream backend integrations will determine whether M5-specific W8A8 kernels reproducibly improve LLM prefill throughput by roughly 1.4x without material accuracy loss.
expiredknownscott: low
Independent benchmarks will determine whether CachyLlama’s SSD-backed multi-tier persistent KV cache materially reduces repeated prompt-processing latency in long local-agent sessions on slower hardware without unacceptable storage or correctness tradeoffs.
expirednovelscott: none
Independent benchmarks will determine whether BeeLlama.cpp's KVarN and low-bit KV-cache formats substantially reduce long-context VRAM use while preserving quality and practical inference speed.
expirednovelscott: low
Independent testing will determine whether llama.cpp’s merged MiniMax-M3 vision support enables reliable local multimodal inference across commonly used hardware.
expirednovelscott: none
llama.cpp maintainers will revert or gate default loading of bundled MTP tensors after reports that models consume extra RAM or VRAM even when MTP speculative decoding is disabled.
expirednovelscott: low
Independent benchmarks will determine whether sampler-level self-signaling of reasoning budgets in llama.cpp improves low-budget coding performance over hard token cutoffs without materially hurting quality at larger budgets.
expired
Independent benchmarks will determine whether llama.cpp’s hot-expert GPU cache materially accelerates CPU-offloaded MoE inference on memory-constrained GPUs without regressions across models and quantizations.
expiredconvergesscott: medium
Independent use will determine whether llama.cpp’s merged Qwen3-TTS support enables reliable, practical local multilingual voice cloning from reference audio.
expiredconvergesscott: high
Independent benchmarks will determine whether FerroX matches llama.cpp in GGUF model compatibility and practical inference performance.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s x86 VNNI Q2_0 kernel delivers roughly 3–3.6x faster CPU inference across representative models without quality or compatibility regressions.
expiredconvergesscott: medium
Independent benchmarks will determine whether llama.cpp's SYCL TILE-kernel dispatch materially accelerates long-context quantized-KV decoding on Intel Battlemage GPUs across representative models and contexts.
expiredknownscott: low
llama.cpp will merge PR 26291, and broader testing will determine whether its configurable RPC loading threads substantially reduce very-large-model load times across distributed hardware without serving regressions.
expiredknownscott: low
llama.cpp will merge LongCat-Flash support, and broader testing will confirm that larger LongCat-Flash GGUF variants run locally without major correctness or compatibility failures.
expiredknownscott: medium
Upstream review and independent benchmarks will determine whether correcting llama.cpp’s inflated MTP buffer reservations materially expands usable context on memory-constrained AMD systems without inference regressions.
resolvedknownscott: low
Independent Linux and Windows testing will determine whether llama.cpp’s ROCm 7.14 targets make AMD’s TheRock-based production stack reliable for local inference.
resolvedconvergesscott: medium
llama.cpp will merge Kimi K3 support, and independent testing will determine whether it enables correct and practical local inference across common hardware configurations.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s adaptive MTP mode selects speculative-decoding depth effectively enough to improve coding-agent throughput without manual tuning.
expiredknownscott: medium
Independent testing will determine whether llama.cpp’s merged Bonsai and ternary-model support enables correct, performant local inference for 1-bit and 1.58-bit Bonsai checkpoints across common backends.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s proposed dense-model CPU-FFN offload materially reduces VRAM requirements while preserving practical throughput for large quantized models.
expiredconvergesscott: medium
Independent benchmarks and merge review will determine whether llama.cpp’s proposed AVX2 IQ kernels materially accelerate large-batch CPU prompt processing without meaningful perplexity loss.
expiredknownscott: medium
Independent testing will determine whether llama.cpp’s dots3-note integration enables correct and practically useful local multimodal inference for the 280B-total, 16B-active model at long context lengths.
expired
Independent benchmarks will determine whether AMD-Ecosystem’s maintained llama.cpp branch materially accelerates ROCm prompt processing on AMD integrated GPUs without unacceptable decode or compatibility tradeoffs.
expiredknownscott: low
Independent use will determine whether Picchio reliably exposes llama.cpp layer placement and separates prefill from decode performance well enough to prevent misleading local-inference benchmarks.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s GLM-4.5-Air MTP support materially improves local inference speed across memory-rich, compute-limited hardware without reducing output quality or stability.
expiredknownscott: medium
Independent benchmarks will determine whether ConvRot’s llama.cpp-compatible Q5 and Q6 quantizations preserve near-Q8 model quality at comparable low-bit memory use without unacceptable performance or stability tradeoffs.
expiredknownscott: medium
Independent testing will determine whether aikitoria’s open NVIDIA kernel fork reliably enables peer-to-peer transfers on supported dual-consumer-GPU systems and materially improves local inference performance.
expiredknownscott: medium
llama.cpp contributor ngxson claims lazy tensor loading can avoid loading unused tensors from large sparse Qwen-family models, materially reducing the RAM or VRAM needed for local inference.
expiredknownscott: medium
Daxfortuna reports that llama.cpp quantization fallbacks leave some GGUF files labeled as lower-bit formats than their tensors actually use, affecting 64 of 443 audited files and undermining reproducible local-model packaging.
expiredconvergesscott: medium
The fork author claims selectively placing frequently used MoE experts in VRAM raises llama.cpp generation throughput from 20 to 30 tokens per second on partially offloaded coding workloads, potentially improving local inference on memory-constrained GPUs.
resolvedknownscott: medium
mattescala claims a llama.cpp NUMA weight-mirroring implementation improves dual-socket CPU decode throughput by 64–137% by replicating weights per NUMA node, trading doubled weight memory for materially better local-inference performance.
expiredconvergesscott: medium
llama.cpp contributor ynankani claims the proposed CUDA MoE fusion for speculative decoding materially accelerates multi-token prediction across sparse models, potentially improving local draft-token throughput if merged.
watchingknownscott: medium
llama.cpp contributor bartowski1182 claims PR #27402 materially accelerates large-batch CPU prompt processing for IQ-quantized models on AVX2 hardware, potentially improving CPU inference throughput if merged.
expiredknownscott: medium
llama.cpp contributor predatar claims PR #28086 raises IQ3-quantized MoE decode throughput on Apple Silicon Metal from about 65.6 to 73.9 tokens per second, potentially improving local sparse-model inference if merged.
watchingknownscott: low
gfx906-llama-cpp maintainer milpster claims the updated fork improves prefill throughput by 14–23% and token generation by about 11% over upstream in reported benchmarks, potentially extending the practical usefulness of legacy AMD GCN hardware for local inference.
expirednovelscott: low
llama.cpp contributor pwilkin claims the merged Flash Attention tuning in PR #28102 materially accelerates long-context prefill on AMD RDNA4 hardware, potentially improving local inference responsiveness without a comparable decode-speed gain.
corroboratednovelscott: low
llama.cpp contributor thelittlefireman claims the merged GCN-specific MMQ configuration improves prompt processing by about 5% in a published gfx906 benchmark, potentially accelerating local inference on older AMD MI50/MI60-class hardware.
watchingnovelscott: low
llamAmpere’s creator claims the released Ampere-focused llama.cpp fork sustains over 90 tokens per second through 100K tokens of context in its recommended coding configuration, potentially making long-context local agents more responsive on RTX 3090-class hardware.
watchingconvergesscott: medium
Luigi reports that Vulkan runs a quantized Qwen3.6-35B MoE at 32.75 generation tokens per second on a Panther Lake laptop, outperforming CPU and the tested SYCL configuration while OpenVINO fails, making backend choice consequential for this local-inference setup.
watchingknownscott: low
Redditor deathcom65 reports that nasone32's specialized llama.cpp fork raises Qwen3.8 Q8 decode throughput from about 28 to 82 tokens per second at 60K context on dual Radeon 7900 XTX GPUs, potentially making long-context local agents substantially more responsive on consumer AMD hardware.
watchingknownscott: low
Fork author neuralll claims his released llama.cpp fork's VRAM-filling hot-expert cache (built on csantiago78's PR #27861) roughly doubles decode throughput for GLM-5.3-Flash and MiMo MoE models far larger than total VRAM on two RTX 3090s with unchanged perplexity, and projects further gains per added GPU — if independent multi-GPU users reproduce it, hot-expert caching becomes a practical standard path for memory-constrained local MoE inference.
acceleratingconvergesscott: high
jbooth's merged llama.cpp PR #27851 claims a tiled VNNI mul_mat path accelerates CPU k-quant prompt processing 3-7x on x86 (about 2x over repack) with microscopic error, and confirmation of the gains on broader hardware plus shipping in releases would make tiled CPU prefill a standard optimization for CPU-served local inference.
watchingconvergesscott: high
Redditor Odd_Cauliflower_8004's llama.cpp fork swaps byte-identical repeated messages in llama-server's chat parser for one-line references, losslessly cutting agent-loop context and prefill costs; upstream merge or independent adoption would establish dedup-by-reference as a standard optimization for repeated agent-loop content.
seedconvergesscott: medium
Hayder Tirmazi claims four implementation changes to llama.cpp's n-gram caches (unnecessary map copies removed, flat hash maps, and related fixes) make prompt-lookup drafting up to 42x faster with up to 2.6x less memory at unchanged acceptance rates; an upstream llama.cpp merge or independent reproduction would establish lean n-gram-cache engineering as a standard local-inference optimization.
watchingconvergesscott: medium
Infermeld's maintainer (do_u_think_im_spooky) claims the released experimental v0.1.0 Linux kit lets a single GGUF run jointly across an AMD (Vulkan) and NVIDIA (CUDA) GPU under llama.cpp, and independent reproduction by mixed-vendor-GPU owners would establish cross-vendor consumer-GPU inference as a practical local setup.
seedknownscott: low
llama.cpp contributor pratiknarola-t's merged PR #29869 claims few-row MMA Metal mat-mul and batched-copy kernels turn DFlash2 speculative decoding of Qwen3.8-27B from slower than serial decoding (~30 tok/s) into ~110 tok/s on an M3 Ultra — with the kernels, tests, and benchmarks disclosed as Claude Code-written — and the gains replicating across Apple GPUs, models, and specdec drafters would establish few-row matmul optimization as the standard enabler of speculative decoding on non-tensor-API Apple Silicon.
seedconvergesscott: high
am17an's merged llama.cpp PR #26610 adds RPC '-sm tensor' — tensor parallelism across networked machines (author-demonstrated on 2x DGX Sparks over RDMA, independently confirmed by ryan5rdx on 2x M3 Ultra, running a 284B-param MoE at 619 pp2048 / ~20 tg128) — and becomes a practical multi-machine local-inference pattern if outside users adopt it across their own clusters with replicated throughput; quiet disuse after the merge closes it.
watchingconvergesscott: high
Microsoft's October 7 keynote and companion first-party blogs claim local LLM inference is now a first-class Windows path — Windows ML shipping experimental llama.cpp/GGUF support today, DeepSeek V4 Flash running locally in 60GB on RTX Spark, and GitHub Copilot gaining local models (MAI Code 1.1 Flash at ~70.8% SWE-Bench Verified on-device) with MXC-sandboxed tool execution by end of October — and on-schedule Copilot shipping plus real developer adoption of the Windows ML stack confirms local inference as mainstream on Windows, while slippage or quiet fade refutes it.
corroboratedconvergesscott: high

Trajectory notes