2026-10-11 17:11 UTC

Independent benchmarks will determine whether BeeLlama.cpp's KVarN and low-bit KV-cache formats substantially reduce long-context VRAM use while preserving quality and practical inference speed.

state: expiredheat: lowuncertainty: highnovelscott: lowkv-cache-quantization llama-cpp local-inferenceAnbeeldBeeLlama.cpp

What is this?

BeeLlama.cpp is Anbeeld’s performance-focused fork of llama.cpp for local GGUF inference, adding variance-normalized KVarN KV-cache quantization, independently selectable K/V bit widths, low-bit cache formats, and a high-precision tail for recent tokens. Anbeeld’s own limited benchmarks on Qwen 3.6 27B report favorable quality-versus-VRAM results, including stronger low-bit behavior when evaluated with KLD, but the author explicitly describes the tests as narrow and “not the whole truth.” The supplied snippets do not establish independent validation across models, context lengths, hardware, output quality, or practical inference speed, so the broader performance claim remains unconfirmed.

Why it matters to Scott

No intersection found in Scott’s wikis, and no radar pages already track BeeLlama.cpp, KVarN, or this benchmark claim. The case may fit Scott’s general local-inference territory, but the supplied hits do not establish a specific position, project impact, or actionable connection.
queries asked of Scott's wikis
  • KV-cache quantization quality and long-context memory tradeoffs
  • local inference VRAM economics on consumer hardware
  • llama.cpp forks versus upstream adoption strategy
  • benchmark standards for quantized inference quality
  • high-precision recent-token tails in agent memory
  • long-context local models for coding-agent workloads

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (9) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditBeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM
LocalLLaMA
Anbeeld2014
🟧 echo.blog ⭐Primary benchmark source. It reports that a 1024-token exact KV tail sharply improves low-bit KLD; its recommendation table gives `kvarn5 + Anbeeld——
🟠 redditKV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates
LocalLLaMA
Anbeeld8833
🟧 hnINT2 KV-cache quantization is getting surprisingly goodyunuyean10
🟠 redditBeeLLama issues
LocalLLaMA
uber-linny17
🟠 reddit1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"
LocalLLaMA
Anbeeld7739
🟠 redditGemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show
LocalLLaMA
Anbeeld6819
🟠 redditQwen 3.8 27B KV f16 vs q8_0 are not equivalents
LocalLLaMA
Felixls51131
🟠 redditTesla P40 - use F16 KV instead of Q8
LocalLLaMA
PairOfRussels15

Interpretation history

Decision trace