2026-10-11 16:38 UTC

Reddit builder Mahmoud (ghraibeh on GitHub) claims his released MIT kNN cache — local bge-small embeddings answering when the nearest stored input is ≥0.90 similar and 5 neighbours agree, CPU-only — delivers author-measured 87% warm-call savings at 97.6% local-answer accuracy in front of Jev-class judgments, and independent adoption or replication of those numbers makes local semantic caching a standard cost-reduction layer for agent decision calls, while a hobby-demo fade closes it.

state: seedheat: lowuncertainty: mediumcontradictsscott: highsemantic-caching inference-economics agent-harnessesghraibeh (Mahmoud)

What is this?

Mahmoud (GitHub user ghraibeh) has released an MIT-licensed local kNN semantic cache: it embeds incoming calls with local bge-small embeddings on CPU only, returns a stored answer when the nearest cached input scores ≥0.90 similarity AND 5 neighbours agree, and posts author-measured 87% warm-call savings at 97.6% local-answer accuracy in front of agent judgment calls. The supplied web material corroborates the surrounding pattern but not the project itself — nothing surfaced the Reddit post, the repo, or any independent replication, so the headline numbers remain author-only. Adjacent hits show semantic caching is an established cost-reduction layer (hobby-to-production projects claiming 20–85% savings at typical 0.90–0.95 thresholds; bge-small-en-v1.5 is the community's default local embedding model), though no supplied snippet matches Mahmoud's 5-neighbour consensus gate. Notably, the closest industry guidance — a 2026 multi-agent caching paper's threshold table — recommends against caching decision/safety-type calls because small input differences flip correct answers, which bears directly on whether a cache in front of judgment calls is sound.

Why it matters to Scott

Author-measured 97.6% accuracy for replaying agent judgment calls from a ≥0.90/5-neighbour kNN gate presses on a seam Scott's canon draws deliberately: the scout–senior split keeps judging the governed, un-cached end, and advisory-embedding-recall holds that embedding similarity is a hint that is never citable — so a consensus gate making similarity itself the warrant is a live challenge, where replication would collapse the two-speed split and a hobby fade would be a dated receipt for it. The challenge is unlanded (grounding found nothing beyond the author's own numbers, and a 2026 caching paper recommends against caching decision calls), but the artifact is first-party and quantified, with lineage in the Redis LangCache episode and the Coalent invalidation problem — the exact correctness question a judgment cache must answer.
ip:source.the-scout-and-the-senior-ebookdev:concept.advisory-embedding-recalldev:concept.web-research-librarydev:concept.cheap-model-front-doorradar:redis-langcache-inference-savingsradar:coalent-source-aware-cache-invalidationradar:concept.llm-cachingradar:concept.inference-economicsradar:concept.embeddings
queries asked of Scott's wikis
  • agent harness LLM call cost reduction caching layer
  • local bge-small embeddings RAG memory pipeline
  • semantic cache threshold false-hit accuracy tradeoff
  • inference economics agent loop token spend
  • agent memory dedupe nearest-neighbour reuse
  • CPU-only local inference constraints

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 164h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-04 19:48⭐ origin directly observedI built a local kNN cache in front of Jev. Numbers, caveats and two negative results inside (author here)
Aggressive-East-2815 on r/LocalLLaMA
—
10-04 19:48amplified on r/LocalLLaMA 👑reddit.post.1wxoqqp
Aggressive-East-2815
peak 0 · 1 comments · 109% of case engagement
10-04 20:20our radar first saw it · +0.5hdiscovery anchor: reddit.post.1wxoqqp—
pace: p9 vs 1247 stories at the 96h mark (now 164h old) — behind 3jsbench-llm-3d-generation-benchmark (0.5x)

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I built a local kNN cache in front of Jev. Numbers, caveats and two negative results inside (author here)
LocalLLaMA
Aggressive-East-281501

Interpretation history

Decision trace