focus-llama creator Ok-Shower7286 claims their llama.cpp fork lets models restrict subsequent attention to self-selected context chunks through output tags without training, potentially reducing long-context decoding costs.
state: seedheat: lowuncertainty: mediumconvergesscott: mediumlong-context-inference attention-optimization local-inferenceOk-Shower7286
What is this?
The case identifies focus-llama as a llama.cpp fork by the pseudonymous creator Ok-Shower7286, who claims it implements “Declarative Attention”: model-generated tags select which context chunks subsequent attention can access, without additional training. The supplied web snippets establish llama.cpp as an open-source C/C++ inference library, but none directly documents focus-llama, its creator, or the cited paper arXiv:2609.02737. The fork’s implementation, training-free operation, and potential long-context decoding savings therefore remain case-level claims, with no supplied benchmarks establishing speed or quality trade-offs.
Why it matters to Scott
The claimed mechanism converges with Scott’s Context Engineering and agent-authored context compaction, but offers a concrete extension worth testing: model-directed attention masking inside inference rather than harness-side removal or summarization of history. This could change how he implements selective context, but the supplied material does not independently establish the fork, training-free operation, or speed–quality trade-offs; the radar’s related Plurnk addressable-context case does not already cover this development.
ip:framework.context-engineeringip:concept.attention-budgetdev:concept.agent-authored-context-compactionradar:plurnk-addressable-context-harnessradar:concept.llama-cppradar:concept.inference-optimization
queries asked of Scott's wikis
- model-directed context selection and attention control
- local inference latency economics llama.cpp
- long-context agents memory retrieval versus full context
- inference optimization benchmarks quality trade-offs
- output tags controlling runtime agent harness behavior
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 530h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p60 vs 1032 stories at the 336h mark (now 530h old) — ahead of openai-2030-burn-projection (1.0x), behind openai-mentalhealthbench (0.9x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-20T14:26:52Z
grounded: converges/medium — The claimed mechanism converges with Scott’s Context Engineering and agent-authored context compaction, but offers a concrete extension worth testing: model-dir
2026-09-20T14:23:25Z
origin walked (codex/luna, conf 0.97): anchor reddit.post.1wli0wz -> echo.github.f116728a28 by edwardyoon
2026-09-20T14:22:21Z
case created — The creator describes a concrete inference implementation distinct from existing cases, but cites paper speedups rather than measured results from the fork.
Decision trace
- 10-10 20:06review_dormantscheduled targets exhausted or 28 quiet days
- 10-10 20:06drop_targetsquiet through full ladder or over cap 8
- 09-25 12:25review_screenjev screen: no material development (noul=0.04)
- 09-22 10:24review_screenThe comments reiterate the already-known paper-based speedup, lack of fork measurements, and uncertainty about practical implementation; they add no verified new result or consequential change.
- 09-22 10:24review_screenjev screen borderline (noul=0.42) — luna review
- 09-21 13:21sensor_dirtycomment_update
- 09-21 00:26groundThe claimed mechanism converges with Scott’s Context Engineering and agent-authored context compaction, but offers a concrete extension worth testing: model-directed attention masking inside inference
- 09-21 00:23promote_anchororigin walk conf 0.97
- 09-21 00:22createThe creator describes a concrete inference implementation distinct from existing cases, but cites paper speedups rather than measured results from the fork.