2026-10-11 16:38 UTC

focus-llama creator Ok-Shower7286 claims their llama.cpp fork lets models restrict subsequent attention to self-selected context chunks through output tags without training, potentially reducing long-context decoding costs.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumlong-context-inference attention-optimization local-inferenceOk-Shower7286

What is this?

The case identifies focus-llama as a llama.cpp fork by the pseudonymous creator Ok-Shower7286, who claims it implements “Declarative Attention”: model-generated tags select which context chunks subsequent attention can access, without additional training. The supplied web snippets establish llama.cpp as an open-source C/C++ inference library, but none directly documents focus-llama, its creator, or the cited paper arXiv:2609.02737. The fork’s implementation, training-free operation, and potential long-context decoding savings therefore remain case-level claims, with no supplied benchmarks establishing speed or quality trade-offs.

Why it matters to Scott

The claimed mechanism converges with Scott’s Context Engineering and agent-authored context compaction, but offers a concrete extension worth testing: model-directed attention masking inside inference rather than harness-side removal or summarization of history. This could change how he implements selective context, but the supplied material does not independently establish the fork, training-free operation, or speed–quality trade-offs; the radar’s related Plurnk addressable-context case does not already cover this development.
ip:framework.context-engineeringip:concept.attention-budgetdev:concept.agent-authored-context-compactionradar:plurnk-addressable-context-harnessradar:concept.llama-cppradar:concept.inference-optimization
queries asked of Scott's wikis
  • model-directed context selection and attention control
  • local inference latency economics llama.cpp
  • long-context agents memory retrieval versus full context
  • inference optimization benchmarks quality trade-offs
  • output tags controlling runtime agent harness behavior

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 530h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-19 14:00⭐ origin echo-reconstructedThe repository’s README describes focus-llama as “a llama.cpp fork for Dynamic Attention Masking,” implementing Declarative Attention throug
edwardyoon on github (echo) · attributed from reddit.post.1wli0wz
—
09-20 14:07first on r/LocalLLaMA · published · +24.1hfocus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737)
Ok-Shower7286
—
09-20 14:07amplified on r/LocalLLaMA 👑reddit.post.1wli0wz
Ok-Shower7286
peak 38 · 8 comments · 100% of case engagement
09-20 14:20our radar first saw it · +24.3hdiscovery anchor: reddit.post.1wli0wz—
pace: p60 vs 1032 stories at the 336h mark (now 530h old) — ahead of openai-2030-burn-projection (1.0x), behind openai-mentalhealthbench (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditfocus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737)
LocalLLaMA
Ok-Shower7286388
🟧 echo.github ⭐The repository’s README describes focus-llama as “a llama.cpp fork for Dynamic Attention Masking,” implementing Declarative Attention througedwardyoon——

Interpretation history

Decision trace