Reddit builder Mahmoud (ghraibeh on GitHub) claims his released MIT kNN cache — local bge-small embeddings answering when the nearest stored input is ≥0.90 similar and 5 neighbours agree, CPU-only — delivers author-measured 87% warm-call savings at 97.6% local-answer accuracy in front of Jev-class judgments, and independent adoption or replication of those numbers makes local semantic caching a standard cost-reduction layer for agent decision calls, while a hobby-demo fade closes it.
state: seedheat: lowuncertainty: mediumcontradictsscott: highsemantic-caching inference-economics agent-harnessesghraibeh (Mahmoud)
What is this?
Mahmoud (GitHub user ghraibeh) has released an MIT-licensed local kNN semantic cache: it embeds incoming calls with local bge-small embeddings on CPU only, returns a stored answer when the nearest cached input scores ≥0.90 similarity AND 5 neighbours agree, and posts author-measured 87% warm-call savings at 97.6% local-answer accuracy in front of agent judgment calls. The supplied web material corroborates the surrounding pattern but not the project itself — nothing surfaced the Reddit post, the repo, or any independent replication, so the headline numbers remain author-only. Adjacent hits show semantic caching is an established cost-reduction layer (hobby-to-production projects claiming 20–85% savings at typical 0.90–0.95 thresholds; bge-small-en-v1.5 is the community's default local embedding model), though no supplied snippet matches Mahmoud's 5-neighbour consensus gate. Notably, the closest industry guidance — a 2026 multi-agent caching paper's threshold table — recommends against caching decision/safety-type calls because small input differences flip correct answers, which bears directly on whether a cache in front of judgment calls is sound.
Why it matters to Scott
Author-measured 97.6% accuracy for replaying agent judgment calls from a ≥0.90/5-neighbour kNN gate presses on a seam Scott's canon draws deliberately: the scout–senior split keeps judging the governed, un-cached end, and advisory-embedding-recall holds that embedding similarity is a hint that is never citable — so a consensus gate making similarity itself the warrant is a live challenge, where replication would collapse the two-speed split and a hobby fade would be a dated receipt for it. The challenge is unlanded (grounding found nothing beyond the author's own numbers, and a 2026 caching paper recommends against caching decision calls), but the artifact is first-party and quantified, with lineage in the Redis LangCache episode and the Coalent invalidation problem — the exact correctness question a judgment cache must answer.
ip:source.the-scout-and-the-senior-ebookdev:concept.advisory-embedding-recalldev:concept.web-research-librarydev:concept.cheap-model-front-doorradar:redis-langcache-inference-savingsradar:coalent-source-aware-cache-invalidationradar:concept.llm-cachingradar:concept.inference-economicsradar:concept.embeddings
queries asked of Scott's wikis
- agent harness LLM call cost reduction caching layer
- local bge-small embeddings RAG memory pipeline
- semantic cache threshold false-hit accuracy tradeoff
- inference economics agent loop token spend
- agent memory dedupe nearest-neighbour reuse
- CPU-only local inference constraints
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 164h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p9 vs 1247 stories at the 96h mark (now 164h old) — behind 3jsbench-llm-3d-generation-benchmark (0.5x)
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-10-04T20:45:50Z
grounded: contradicts/high — Author-measured 97.6% accuracy for replaying agent judgment calls from a ≥0.90/5-neighbour kNN gate presses on a seam Scott's canon draws deliberately: the scou
2026-10-04T20:37:08Z
case created — First-party released artifact with quantified savings and unusually honest negative caveats — a usable concrete event per the radar's artifact rule, despite near-zero engagement.
Decision trace
- 10-05 07:45groundAuthor-measured 97.6% accuracy for replaying agent judgment calls from a ≥0.90/5-neighbour kNN gate presses on a seam Scott's canon draws deliberately: the scout–senior split keeps judging the go
- 10-05 07:37createFirst-party released artifact with quantified savings and unusually honest negative caveats — a usable concrete event per the radar's artifact rule, despite near-zero engagement.