2026-10-11 16:37 UTC

Yandex Research claims manipulating an LLM’s KV-cache as agent runtime state can improve interactivity and responsiveness, potentially enabling agents to handle changing inputs without conventional turn-by-turn inference.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumkv-cache-agent-runtime interactive-inference agent-harnessesYandex Research

What is this?

Yandex Research published a position piece, 'The KV cache as an agent runtime' (research.yandex.com, cross-posted to Yandex's Medium), arguing that interactivity is largely an inference-runtime problem: by changing how the KV cache — the transformer's reusable execution state — is partitioned, ordered, exposed, and scheduled, an inference engine can implement interaction protocols absent from training, framing multimodal agents as asynchronous I/O systems; the post includes a 'training-free Doom agent' demo section and builds on the team's earlier Hogwild! Inference and AsyncReasoning work, though the snippets show no measured responsiveness results and no named authors. The snippets show the broader direction was already active before the blog: an arXiv paper (2603.04428, ~March 2026) demonstrates persistent quantized KV cache as working memory across multi-agent phases with large time-to-first-token reductions, an ICLR 2026 workshop paper (Continuum) proposes KV-cache TTL scheduling for multi-turn agent serving, and AMD ships KV-cache reuse/rewind in its Ryzen AI local-inference stack. None of these demonstrates Yandex's specific claim — interaction protocols or changing-input handling without turn-by-turn inference — so that remains untested on the supplied material, while 'KV cache as manipulable agent runtime state' is corroborated as a live research and productization direction across academia, serving stacks, and local inference.

Why it matters to Scott

Yandex's move — interaction protocols implemented in the inference runtime rather than the chat turn — independently arrives at the core of his Agent-Native Computing substrate critique and his Model-Plus-Harness claim that capability lives in the runtime, and the corroboration sweep (practitioner KV-cache transplants on a 32GB consumer GPU, Cache-to-Cache model-to-model KV transfer, AMD shipping cache reuse in a local stack) makes cache-level state manipulation practiced on exactly his local-inference turf. But Yandex's specific claim — responsiveness without turn-by-turn inference — still has no implementation or measurement, so today's value is dated receipts for his positions plus a watch on whether KV-level compaction/transfer converges with his agent-authored compaction and provider-bound-reasoning-continuity concepts.
ip:framework.agent-native-computingip:concept.model-plus-harness-benchmark-unitip:framework.context-engineeringdev:concept.hardware-aware-local-inferencedev:concept.provider-bound-reasoning-continuityradar:concept.kv-cacheradar:concept.agent-harnessesradar:concept.local-inferenceradar:spomin-live-kv-compactionradar:zero-copy-kv-cache-migratorradar:cachyllama-persistent-kv-cache
queries asked of Scott's wikis
  • Agent-Native Computing chat-turn substrate critique
  • Model-Plus-Harness Bench runtime evaluation criteria
  • agent memory working state vs in-context KV cache
  • local inference stack latency interactivity llama.cpp
  • cache-friendly append-only context design agent harness
  • training-free agent control loop demo position

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 823h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-07 09:22 (minted)⭐ origin echo-reconstructedThe team's Reddit post describes modifying model inference state (KV-cache) to achieve more interactive and responsive LLM systems, building
Yandex Research on blog (echo) · attributed from reddit.post.1w9myqc · published time unknown
—
09-07 09:03first on r/MachineLearning · published · lag ?KV cache as an agent runtime [R]
_puhsu
—
09-25 20:35first on r/LocalLLaMA · published · lag ?Qwen3.8-27B: Using KV Cache Transplants to Boost Output Quality
wadeAlexC
—
09-07 09:03amplified on r/MachineLearningreddit.post.1w9myqc
_puhsu
peak 15 · 4 comments · 17% of case engagement
09-25 20:35amplified on r/LocalLLaMA 👑reddit.post.1wq76f6
wadeAlexC
peak 68 · 24 comments · 83% of case engagement
09-07 09:20our radar first saw it · lag ?discovery anchor: reddit.post.1w9myqc—
pace: p69 vs 519 stories at the 720h mark (now 823h old) — ahead of dual-dgx-spark-deepseek-flash-v4 (1.0x), behind halv-coding-token-savings (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditKV cache as an agent runtime [R]
MachineLearning
_puhsu154
🟧 echo.blog ⭐The team's Reddit post describes modifying model inference state (KV-cache) to achieve more interactive and responsive LLM systems, buildingYandex Research——
🟠 redditQwen3.8-27B: Using KV Cache Transplants to Boost Output Quality
LocalLLaMA
wadeAlexC6824

Interpretation history

Decision trace