2026-10-11 17:20 UTC

Independent benchmarks will determine whether ExANS can sustain near-622 GB/s lossless BF16 KV-cache compression on H100-class GPUs and materially reduce offload bandwidth and time-to-first-token in long-context serving.

state: expiredheat: lowuncertainty: highconvergesscott: mediumkv-cache llm-inference local-inference memory-bandwidthOpenLake

What is this?

OpenLake reports ExANS, an on-GPU method for losslessly compressing BF16 KV-cache blocks, with claimed throughput of 622 GB/s on an NVIDIA H100. It is intended to reduce the PCIe or network bandwidth required when offloading KV caches and thereby improve time-to-first-token for long-context LLM serving. The supplied results establish KV-cache capacity and movement as important inference bottlenecks, but they do not independently benchmark ExANS or verify its claimed throughput and end-to-end latency benefits.

Why it matters to Scott

ExANS converges with Scott’s hardware-aware local-inference position by treating KV-cache compression, accelerator memory, and offload bandwidth as explicit serving-runtime concerns. If independently validated end to end, it could affect long-context serving architecture and TTFT, but current evidence is only an H100-class performance claim and does not establish benefits for Scott’s consumer/local GPU stack.
dev:concept.hardware-aware-local-inferenceradar:dkv-kv-cache-compression-validationradar:concept.long-context-inferenceradar:concept.inference-efficiency
queries asked of Scott's wikis
  • lossless KV-cache compression vs quantization
  • KV-cache offload and long-context serving
  • memory bandwidth as the inference bottleneck
  • time-to-first-token optimization
  • local inference memory hierarchy
  • benchmarking claims vs end-to-end serving gains

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: ExANS – Lossless KV cache compression at 622 GB/s on H100arnav__1130
🟧 echo.blog ⭐OpenLake reports an on-GPU lossless compression method for BF16 KV-cache blocks reaching 622 GB/s on an H100, intended to reduce PCIe or netOpenLake——

Interpretation history

Decision trace