2026-10-11 17:11 UTC

Independent production evaluations will determine whether Thompson-sampling model routing lowers LLM serving costs versus static routing policies without materially reducing response quality.

state: expiredheat: lowuncertainty: highnovelscott: lowllm-routing online-learning inference-costsRamp

What is this?

Ramp deployed an LLM-routing system that treats model selection as an online-learning problem, using Thompson sampling to adapt routing decisions rather than relying on static configuration. A ZenML case-study snippet reports more than 25% cost savings in a broader experiment, while other supplied sources stress that routing must be evaluated against explicit quality constraints and production-specific traffic. The snippets do not establish that independent production evaluations have yet replicated Ramp’s result or conclusively shown no material quality loss.

Why it matters to Scott

No intersection found in Scott’s wikis, and no radar pages show this development or actor is already tracked. The case is broadly relevant to LLM tooling and inference costs, but the supplied hits do not establish a connection to Scott’s specific positions, projects, or prior coverage.
queries asked of Scott's wikis
  • adaptive model routing versus static policies
  • LLM gateway and centralized inference infrastructure
  • cost-quality evals for model routing
  • online learning and bandits in production systems
  • multi-model inference economics
  • shadow evaluation and per-route quality SLAs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOnline Learning for Cost-Efficient LLM Routinggmays10
🟧 echo.blog ⭐Ramp describes using online learning and Thompson sampling to route LLM requests for lower cost while preserving quality.Ramp——
🟧 hnLaunch HN: Tokenless (YC S26) – Automatic model switching to save moneyrohaga7163
🟧 hnWe built an AI model router that cut LLM costs by 94%inesbarros111

Interpretation history

Decision trace