2026-10-11 16:38 UTC

Redis presents LangCache as reducing repeated LLM inference, citing Mangoes.ai's reported 70% cache hit rate, 70% LLM-spend savings, and fourfold speedup, potentially making caching a material serving-cost control for repetitive application workloads.

state: seedheat: lowuncertainty: mediumconvergesscott: mediuminference-economics semantic-caching llm-toolingRedisMangoes.aiAmit Lamba

What is this?

Redis LangCache is a fully managed semantic-caching service accessed through a REST API that reuses stored LLM responses for similar queries rather than making another model call. Redis's product page and customer story quote Mangoes.ai founder and CEO Amit Lamba reporting a 70% cache hit rate, 70% savings on LLM spend, and fourfold faster responses for its patient-care voice assistant. The customer story attributes the high hit rate to repeated intents and says filtering excludes patient-specific or PII-heavy content from the cache. These are vendor-published customer claims, repeated across the supplied results rather than independently verified; the snippets do not establish measurement methodology, net savings after caching costs, or applicability to other workloads.

Why it matters to Scott

Redis's vendor-published customer account converges with Scott's Derived-Page Cache position on reusing generated answers, and its patient-care voice workload offers a concrete caching experiment for the DentalQuest receptionist offer he helped build—not merely another cost-saving example. The reported gains justify testing repeated, non-patient-specific intents, but do not establish net savings, answer correctness or source-change invalidation; this is response reuse, not validation of his distinct Context Arbitrage mechanism.
ip:concept.derived-page-cachework:concept.dynaquest-ai-dental-receptionistwork:project.marife-ramirezradar:concept.llm-cachingradar:concept.inference-economicsradar:coalent-source-aware-cache-invalidation
queries asked of Scott's wikis
  • LLM serving economics response caching versus repeated inference
  • Semantic caching similarity thresholds answer correctness evaluation
  • Agent memory reusable answers cache invalidation freshness
  • RAG caching privacy patient-specific context filtering
  • Managed LLM infrastructure versus custom caching layers

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 727h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 09:26 (minted)⭐ origin echo-reconstructedRedis's LangCache page quotes Mangoes.ai founder Amit Lamba reporting a 70% cache hit rate, 70% savings on LLM spend, and a fourfold speedup
Redis on blog (echo) · attributed from hn.story.49655370 · published time unknown
—
09-11 08:54first on hacker news · published · lag ?Redis built a cache that cuts LLM costs
soltanov
—
09-11 08:54amplified on hacker news 👑hn.story.49655370
soltanov
peak 1 · 0 comments · 106% of case engagement
09-11 09:21our radar first saw it · lag ?discovery anchor: hn.story.49655370—
pace: p11 vs 519 stories at the 720h mark (now 727h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnRedis built a cache that cuts LLM costs
Retrieved article excerpt

Open article · Retrieved 2026-09-11T09:23:05.991471+00:00

" "Our voice app for patient care gets a lot of specific treatment questions, so it has to be absolutely accurate, and that's what LangCache does. I was worried about LLM costs for high usage, but with LangCache, we're getting a 70% cache hit rate, which saves 70% of our LLM spend. On top of that, it’s 4X faster, which makes a huge difference for real-time patient interactions." Amit Lamba Founder & CEO, Mangoes.ai
soltanov10
🟧 echo.blog ⭐Redis's LangCache page quotes Mangoes.ai founder Amit Lamba reporting a 70% cache hit rate, 70% savings on LLM spend, and a fourfold speedupRedis——

Interpretation history

Decision trace