2026-10-11 17:10 UTC

Complex KDAโ€™s authors claim their released architecture, code, and checkpoints improve the expressivity of Kimi Delta Attention while retaining efficient recurrent long-context execution, potentially making delta-rule models a more practical transformer alternative.

state: seedheat: lowuncertainty: mediumknownscott: mediumalternative-architectures linear-rnn local-inferenceOpenEuroLLMJulien SiemsRiccardo GrazziAntonio OrvietoAaron Klein

What is this?

Complex KDA (CKDA) is a research paper (arXiv:2609.24797, submitted 21 Sep 2026) by Julien Siems, Riccardo Grazzi, Antonio Orvieto, Aaron Klein and colleagues (including Frank Hutter and Jenia Jitsev). It modifies Kimi Delta Attention โ€” the linear-attention/delta-rule mechanism behind Moonshot AI's Kimi Linear hybrid architecture, which open-sourced its kernel, vLLM integration, and checkpoints โ€” by widening parameter ranges (gates to [-1,1], delta-rule coefficient ฮฒ to [0,2]) so a single delta-rule transition plus the channel-wise gate can realize 2D rotations. The claimed result: state-tracking expressivity matching DeltaProduct_2 while keeping KDA's diagonal-plus-rank-one, non-expansive, efficiently recurrent structure. The snippets support the theory and the CKDA formulation, but do not themselves show trained CKDA checkpoints or head-to-head quality results; the checkpoint releases cited in the snippets belong to the original Kimi Linear work, and the hypothesis's mention of OpenEuroLLM as an author affiliation is not confirmed by the supplied material.

Why it matters to Scott

The radar already tracks this exact story in radar:kimi-linear-local-validation โ€” whether Kimi Linear's delta-rule hybrid is a practical local alternative to transformers โ€” and CKDA is a theory-side development on that same architecture, attacking the state-tracking expressivity limit that is the standard objection to linear attention. It matters because a confirmed result (and trained CKDA checkpoints, which the supplied snippets do NOT show โ€” only the original Kimi Linear releases do) would directly bear on whether constant-memory delta-rule models earn a slot on Scott's gamepc/Ollama local-inference stack; as it stands it's a credible theory claim extending an open case, not new information about Scott's own positions.
dev:concept.hardware-aware-local-inferencedev:technology.ollamaradar:kimi-linear-local-validationradar:concept.alternative-architecturesradar:concept.recurrent-modelsradar:concept.model-architecture
queries asked of Scott's wikis
  • delta-rule linear RNN vs transformer architecture argument
  • constant-memory state models for local inference cost
  • agent memory as fixed-size recurrent state / fast weights
  • state tracking and expressivity limits of linear attention
  • open checkpoints of non-transformer architectures for on-device runs
  • hybrid attention ratios in long-context harness design

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 465h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-22 06:34โญ origin directly observed[R] Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Yossarian_1234 on r/LocalLLaMA
โ€”
09-22 06:34amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wn1t6a
Yossarian_1234
peak 23 ยท 6 comments ยท 100% of case engagement
09-22 07:20our radar first saw it ยท +0.8hdiscovery anchor: reddit.post.1wn1t6aโ€”
pace: p57 vs 1032 stories at the 336h mark (now 465h old) โ€” ahead of ai-agent-ransomware-operation (1.0x), behind agent-iap-credential-brokering (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญ[R] Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
LocalLLaMA
Yossarian_1234236

Interpretation history

Decision trace