Yandex Research claims manipulating an LLM’s KV-cache as agent runtime state can improve interactivity and responsiveness, potentially enabling agents to handle changing inputs without conventional turn-by-turn inference.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumkv-cache-agent-runtime interactive-inference agent-harnessesYandex Research
What is this?
Yandex Research published a position piece, 'The KV cache as an agent runtime' (research.yandex.com, cross-posted to Yandex's Medium), arguing that interactivity is largely an inference-runtime problem: by changing how the KV cache — the transformer's reusable execution state — is partitioned, ordered, exposed, and scheduled, an inference engine can implement interaction protocols absent from training, framing multimodal agents as asynchronous I/O systems; the post includes a 'training-free Doom agent' demo section and builds on the team's earlier Hogwild! Inference and AsyncReasoning work, though the snippets show no measured responsiveness results and no named authors. The snippets show the broader direction was already active before the blog: an arXiv paper (2603.04428, ~March 2026) demonstrates persistent quantized KV cache as working memory across multi-agent phases with large time-to-first-token reductions, an ICLR 2026 workshop paper (Continuum) proposes KV-cache TTL scheduling for multi-turn agent serving, and AMD ships KV-cache reuse/rewind in its Ryzen AI local-inference stack. None of these demonstrates Yandex's specific claim — interaction protocols or changing-input handling without turn-by-turn inference — so that remains untested on the supplied material, while 'KV cache as manipulable agent runtime state' is corroborated as a live research and productization direction across academia, serving stacks, and local inference.
Why it matters to Scott
Yandex's move — interaction protocols implemented in the inference runtime rather than the chat turn — independently arrives at the core of his Agent-Native Computing substrate critique and his Model-Plus-Harness claim that capability lives in the runtime, and the corroboration sweep (practitioner KV-cache transplants on a 32GB consumer GPU, Cache-to-Cache model-to-model KV transfer, AMD shipping cache reuse in a local stack) makes cache-level state manipulation practiced on exactly his local-inference turf. But Yandex's specific claim — responsiveness without turn-by-turn inference — still has no implementation or measurement, so today's value is dated receipts for his positions plus a watch on whether KV-level compaction/transfer converges with his agent-authored compaction and provider-bound-reasoning-continuity concepts.
ip:framework.agent-native-computingip:concept.model-plus-harness-benchmark-unitip:framework.context-engineeringdev:concept.hardware-aware-local-inferencedev:concept.provider-bound-reasoning-continuityradar:concept.kv-cacheradar:concept.agent-harnessesradar:concept.local-inferenceradar:spomin-live-kv-compactionradar:zero-copy-kv-cache-migratorradar:cachyllama-persistent-kv-cache
queries asked of Scott's wikis
- Agent-Native Computing chat-turn substrate critique
- Model-Plus-Harness Bench runtime evaluation criteria
- agent memory working state vs in-context KV cache
- local inference stack latency interactivity llama.cpp
- cache-friendly append-only context design agent harness
- training-free agent control loop demo position
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 823h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p69 vs 519 stories at the 720h mark (now 823h old) — ahead of dual-dgx-spark-deepseek-flash-v4 (1.0x), behind halv-coding-token-savings (1.0x)
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-09-26T07:29:06Z
The velocity spike is discussion churn on the already-corroborated transplant thread: new comments are speculative (frontier-lab quantization guess), aspirational (same-cache model switching in dwarfstar), or a request for a tool-use benchmark that hasn't been run — no new implementation, measurement, or validation of Yandex's interactivity claim. Meaning is unchanged from the 09-25 corroboration; momentum is cooling and spread stays single-platform, so heat stays low despite the transient 3.7× peer spike.
2026-09-25T21:51:34Z
grounded: converges/medium — Yandex's move — interaction protocols implemented in the inference runtime rather than the chat turn — independently arrives at the core of his Agent-Native Com
2026-09-25T21:44:41Z
The case's meaning shifts from a single-team architectural proposal to a corroborated research direction: an independent LocalLLaMA practitioner demonstrates KV-cache transplants working on Qwen3.8-27B and cites the Cache-to-Cache paper, making KV-cache manipulation a practiced technique rather than Yandex-only theory. Yandex's specific claim — responsiveness and changing-input handling without turn-by-turn inference — remains untested, so uncertainty only drops to medium and heat stays low despite the hot agent-harnesses neighbourhood (the episode itself runs ~1.2 pts/h, 50th percentile).
2026-09-25T21:25:28Z
evidence attached: reddit.post.1wq76f6 — Practitioner exploration of KV-cache transplants plus the cited Cache-to-Cache paper materially extends the KV-cache-as-agent-runtime line of work with an agent-to-agent context-transfer angle.
2026-09-11T10:28:54Z
This remains a relevant runtime-design proposal rather than an established responsiveness improvement: the staleness check adds no implementation, measurements, or independent validation. Keep it as an architectural reference for Scott, with less frequent review until concrete technical evidence arrives.
2026-09-09T10:24:09Z
No substantive new evidence changes this from an architectural proposal into a demonstrated agent runtime; the team post and reconstructed echo still represent one line of testimony. Its comparison value for Scott remains intact, but there is no supplied implementation or measured responsiveness result to justify promotion.
2026-09-07T09:28:45Z
The supplied reobservation adds no substantive evidence: the team post and reconstructed blog echo remain one line of testimony, not independent validation. KV-cache scheduling remains a relevant architectural comparison for Scott, but measured responsiveness gains and an available implementation are still unestablished.
2026-09-07T09:25:28Z
grounded: converges/medium — Yandex Research’s proposal converges with Scott’s Agent-Native Computing rejection of chat turns as the natural agent substrate and his Model-Plus-Harness Bench
2026-09-07T09:22:49Z
case created — A team-linked research write-up introduces a bounded runtime approach distinct from existing KV-cache capacity and monitoring cases.
Decision trace
- 09-26 17:29repriceThe velocity spike is discussion churn on the already-corroborated transplant thread: new comments are speculative (frontier-lab quantization guess), aspirational (same-cache model switching in dwarfs
- 09-26 14:20sensor_dirtycomment_update
- 09-26 10:21sensor_dirtyvelocity_spike
- 09-26 07:51repriceThe case's meaning shifts from a single-team architectural proposal to a corroborated research direction: an independent LocalLLaMA practitioner demonstrates KV-cache transplants working on Qwen3
- 09-26 07:51groundYandex's move — interaction protocols implemented in the inference runtime rather than the chat turn — independently arrives at the core of his Agent-Native Computing substrate critique and his M
- 09-26 07:44review_reactivatedThe case's meaning shifts from a single-team architectural proposal to a corroborated research direction: an independent LocalLLaMA practitioner demonstrates KV-cache transplants working on Qwen3
- 09-26 07:25attachPractitioner exploration of KV-cache transplants plus the cited Cache-to-Cache paper materially extends the KV-cache-as-agent-runtime line of work with an agent-to-agent context-transfer angle.
- 09-26 07:23propose_attachPractitioner exploration of KV-cache transplants plus the cited Cache-to-Cache paper materially extends the KV-cache-as-agent-runtime line of work with an agent-to-agent context-transfer angle.
- 09-19 04:43review_dormantscheduled targets exhausted or 28 quiet days
- 09-19 04:43drop_targetsquiet through full ladder or over cap 8
- 09-11 20:28repriceThis remains a relevant runtime-design proposal rather than an established responsiveness improvement: the staleness check adds no implementation, measurements, or independent validation. Keep it as a
- 09-11 20:28alert_silentNo consequential new delta is supplied. The team-linked proposal and reconstructed echo remain the same evidence line; neither the adjacent topic's heat nor another scheduled check makes the next
- 09-11 20:28alert_routeNo consequential new delta is supplied. The team-linked proposal and reconstructed echo remain the same evidence line; neither the adjacent topic's heat nor another scheduled check makes the next
- 09-09 20:24repriceNo substantive new evidence changes this from an architectural proposal into a demonstrated agent runtime; the team post and reconstructed echo still represent one line of testimony. Its comparison va
- 09-09 20:24alert_silentThe staleness trigger supplies no consequential new delta. The proposal remains suitable for a regular briefing; neither a newly available runtime nor a concrete result requires Scott’s attention toda
- 09-09 20:24alert_routeThe staleness trigger supplies no consequential new delta. The proposal remains suitable for a regular briefing; neither a newly available runtime nor a concrete result requires Scott’s attention toda
- 09-08 21:21sensor_dirtyengagement_update
- 09-08 03:22sensor_dirtyengagement_update
- 09-07 19:28repriceThe supplied reobservation adds no substantive evidence: the team post and reconstructed blog echo remain one line of testimony, not independent validation. KV-cache scheduling remains a relevant arch
- 09-07 19:28alert_silentNo new release, implementation, or validation changes Scott’s decisions today. The architectural proposal can remain in the next briefing; the engagement change adds no consequential delta.
- 09-07 19:28alert_routeNo new release, implementation, or validation changes Scott’s decisions today. The architectural proposal can remain in the next briefing; the engagement change adds no consequential delta.
- 09-07 19:27alert_silentThe team’s linked architectural write-up and reported DOOM preview offer a substantive comparison for Scott’s agent-native computing work. However, the supplied delta is a synthesis of earlier researc
- 09-07 19:27surface_candidateThe team’s linked architectural write-up and reported DOOM preview offer a substantive comparison for Scott’s agent-native computing work. However, the supplied delta is a synthesis of earlier researc
- 09-07 19:27alert_routeThe team’s linked architectural write-up and reported DOOM preview offer a substantive comparison for Scott’s agent-native computing work. However, the supplied delta is a synthesis of earlier researc
- 09-07 19:25groundYandex Research’s proposal converges with Scott’s Agent-Native Computing rejection of chat turns as the natural agent substrate and his Model-Plus-Harness Benchmark Unit claim that capabilities depend
- 09-07 19:22createA team-linked research write-up introduces a bounded runtime approach distinct from existing KV-cache capacity and monitoring cases.