2026-10-11 16:37 UTC

Redditor Odd_Cauliflower_8004's llama.cpp fork swaps byte-identical repeated messages in llama-server's chat parser for one-line references, losslessly cutting agent-loop context and prefill costs; upstream merge or independent adoption would establish dedup-by-reference as a standard optimization for repeated agent-loop content.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumllama-cpp prompt-deduplication agent-context-costs

What is this?

llama.cpp is the de facto open-source engine for local LLM inference (started by Georgi Gerganov in 2023, MIT-licensed, the core behind Ollama and LM Studio per Wikipedia), and llama-server is its HTTP serving binary. The case rests on a single Reddit post claiming a fork that rewrites byte-identical repeated messages in llama-server's chat handler as one-line references, losslessly shrinking agent-loop prompts and prefill work; the supplied search results do not surface the post or the fork itself, so neither the mechanism's details nor any upstream-merge/adoption status can be verified from them. What the snippets do independently establish is the pain point: a can1357/oh-my-pi issue reports llama-server re-prefilling entire contexts when byte-stability of already-sent messages breaks, confirming that per-turn context inflation and cache invalidation are a live cost problem for agent loops on llama-server. Nothing in the snippets shows a dedup-by-reference mechanism existing anywhere upstream.

Why it matters to Scott

The fork independently instantiates Scott's reference-over-value doctrine โ€” configuration-by-reference-for-agents and url-backed-generative-artifacts already argue against carrying repeated payloads in context โ€” but it arrives at a layer and with a mechanism his canon does not hold: lossless, byte-identical message dedup inside llama-server itself, a lever complementary to (not repeating) his prefix-caching-economics position and directly relevant to his Ollama-on-gamepc serving stack and ask's deliberately lossy --compact path. A single near-zero-engagement unverified post with no merge or adoption evidence keeps this a watch-and-test item rather than something that changes what he argues today; an upstream merge would immediately upgrade it to a dated receipt plus an actionable local-inference optimization.
ip:concept.configuration-by-reference-for-agentsdev:concept.url-backed-generative-artifactsip:concept.prefix-caching-economicsdev:project.askdev:technology.ollamaradar:concept.kv-cacheradar:concept.prompt-cachingradar:concept.token-efficiencyradar:concept.context-managementradar:concept.llama-cppradar:spomin-live-kv-compaction
queries asked of Scott's wikis
  • agent loop context growth token cost
  • KV cache prefix reuse prefill caching
  • context compaction summarization agent memory
  • llama.cpp llama-server local inference serving
  • content-addressed dedup reference compression
  • multi-turn agent harness message history handling

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 410h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-24 14:00โญ origin echo-reconstructedREADME_DEDUPLICATION.md in the fork opens: \"This fork of llama.cpp adds an opt-in pass to `llama-server`. It shortens the prompt when a cha
llopresto87 (GitHub username; evidently the same person as Reddit poster Odd_Cauliflower_8004) on github (echo) ยท attributed from reddit.post.1wqrbdl
โ€”
09-26 14:02first on r/LocalLLaMA ยท published ยท +48.0hI've found a transparent, loseless prompt deduplicator for LLAMA.CPP
Odd_Cauliflower_8004
โ€”
09-26 14:02amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wqrbdl
Odd_Cauliflower_8004
peak 0 ยท 11 comments ยท 99% of case engagement
09-26 14:20our radar first saw it ยท +48.3hdiscovery anchor: reddit.post.1wqrbdlโ€”
pace: p47 vs 1032 stories at the 336h mark (now 410h old) โ€” ahead of anthropic-ci-test-selection-redesign (1.1x), behind blast-sandbox-as-a-service (0.9x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditI've found a transparent, loseless prompt deduplicator for LLAMA.CPP
LocalLLaMA
Odd_Cauliflower_8004011
๐ŸŸง echo.github โญREADME_DEDUPLICATION.md in the fork opens: \"This fork of llama.cpp adds an opt-in pass to `llama-server`. It shortens the prompt when a challopresto87 (GitHub username; evidently the same person as Reddit poster Odd_Cauliflower_8004)โ€”โ€”

Interpretation history

Decision trace