2026-10-11 16:36 UTC

long-context

band: warmmomentum: stable score: 0.463
temperature history

Episodes (13)

Independent evaluations will determine whether Swiss AI’s Apertus 1.5 8B and 70B models combine competitive multilingual quality with practical 262K-context performance while using fully open training data.
expirednovelscott: none
Independent benchmarks will determine whether Meituan’s LongCat-Flash-Lite-Sparse can deliver practical 256K-context inference on 24GB GPUs by combining sparse MoE activation with a RAM-offloaded n-gram lookup table.
expirednovelscott: none
Independent use will determine whether Poolside's updated Laguna S 2.1 FP8/NVFP4 weights fix prior looping failures while reliably supporting the new million-token context.
expirednovelscott: none
Independent reproduction will determine whether DeepSeek V4 Flash can sustain useful million-token inference at practical speeds on a single RTX 5090 using CPU-offloaded experts and adaptive speculative decoding.
resolvedconvergesscott: medium
Independent evaluations will determine whether dots3-note-preview’s 280B-total, 16B-active multimodal MoE architecture and 512K context provide practically competitive quality and efficiency for long-context tool-use and agent workloads.
expiredknownscott: low
Independent benchmarks and artifact review will determine whether the reported sub-2-bit 250M-parameter model can deliver useful local inference from a roughly 60MB deployment while using disk-backed compression for histories approaching 100 million tokens.
expiredknownscott: medium
XHToken claims its released Spark-X2.5 1.7B and 4B models combine native one-million-token context with unusually strong small-model quality, potentially expanding long-context local inference once runtime support matures.
expiredknownscott: medium
Agnes AI claims its released Apache-2.0 Agnes-3.0-Flash supports 262K-context multimodal reasoning and tool use with growing KV caches in only 18 of 72 layers, potentially reducing memory requirements for long-context self-hosted inference.
watchingnovelscott: medium
Early users and Armature leaderboard runs claim Opus 5.5 roughly halves its rebuild-everything behavior in favor of third-party tools and sustains usable quality past 500k tokens of context, marking a real behavioral break from Opus 5 for coding-agent work.
resolvedconvergesscott: high
KnownAd4832 claims a purpose-built single-model inference engine sustains ~65 tok/s decode of Qwen3.8-Flash-Next at 128K context on a 12GB RTX 5070 (~430 tok/s prompt processing, versus ~15 tok/s on llama.cpp), and replication would establish custom model-specific engines as a practical path for low-VRAM long-context local inference.
significantconvergesscott: high
NaiveAI claims its MIT-licensed Naive-N0.5-Flash — a 309B-A15.5B sparse MoE with native 1M-token context via hybrid SWA/DSA and no full-attention layers, served by its AI-optimized NaiveRT stack at up to 2,000 tokens/s — delivers frontier-comparable coding and AI-R&D capability at open weights; independent benchmarking and self-hosted adoption would establish it as a credible local coding model.
watchingnovelscott: medium
Aleph Alpha claims its Apache-2.0 Kolibri-1 — a 78B-parameter MoE with 3.46B active parameters and up to 1M-token context — is a practical compact long-context open model for local and agentic inference; sustained community adoption (quantizations, local deployments, harness integrations) confirms it, while quiet fading after launch marks another release that didn't stick.
resolvedconvergesscott: medium
The authors of arXiv:2610.10845 claim a practical long-term memory mechanism with a 50M token window for LLMs — if validated, it advances agent-memory architectures beyond current context limits.
seedconvergesscott: high

Trajectory notes