2026-10-11 17:11 UTC

Alexia Jolicoeur-Martineau and collaborators claim pretrained LLMs can replace quadratic attention with sliding-window attention and attention sinks without post-training or substantial quality loss, potentially reducing memory requirements for local inference.

state: expiredheat: lowuncertainty: highnovelscott: mediumlocal-inference efficient-attentionAlexia Jolicoeur-Martineau

What is this?

“Sliding Window Attention Adaptation” is a proposed framework for running full-attention-pretrained LLMs with sliding-window attention, using techniques such as persistent attention-sink tokens, interleaved local/global attention, and prefill-only windowing. The supplied snippets report as much as 100% acceleration while largely preserving performance, but also say naïve replacement severely harms long-context performance and that some successful recipes include sliding-window-aware fine-tuning; they therefore do not firmly establish the stronger claim of universally requiring no post-training or causing no substantial quality loss. The snippets also do not establish Alexia Jolicoeur-Martineau’s authorship or identify the collaborators.

Why it matters to Scott

This is not a position already present in Scott’s canon or an exact development already tracked by the radar. If validated, replacing full attention without retraining could materially reduce VRAM requirements on his self-hosted GPU serving stack, but the supplied grounding leaves the no-post-training and quality-preservation claims unresolved.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.long-context-inferenceradar:concept.local-inferenceradar:concept.inference-efficiencyradar:concept.kv-cache
queries asked of Scott's wikis
  • local inference memory and KV-cache economics
  • attention sinks and bounded-context inference
  • training-free adaptation of pretrained models
  • sliding-window versus linear attention
  • long-context quality versus inference efficiency
  • local model serving and efficient attention kernels

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditSliding-window beats linear attention
LocalLLaMA
woadwarrior3620
🟧 echo.paper ⭐The original paper reports: “Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models.Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, and Emy Gervais——
🟠 redditSliding-window attention beats linear on long-context reasoning [R]
MachineLearning
Justgototheeffinmoon277

Interpretation history

Decision trace