2026-10-11 17:12 UTC

Aleph Alpha claims its Apache-2.0 Kolibri-1 โ€” a 78B-parameter MoE with 3.46B active parameters and up to 1M-token context โ€” is a practical compact long-context open model for local and agentic inference; sustained community adoption (quantizations, local deployments, harness integrations) confirms it, while quiet fading after launch marks another release that didn't stick.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumopen-model-release long-context local-inferenceAleph Alpha
Surfaced 2026-10-03T13:38:46Z โ€” Announcing Kolibri-1: 78B parameters, 3.46B active, up to 1M tokens of context, Apache 2.0 โ€” with an X announcement and a published tech rep โ€” The episode crossed from announcement-only to first independent signal: a hands-on local fp8 run (~170 tok/s on a 96GB-class RTX Pro 6000) confirms hardware feasibility and speed while flagging verbose/overthinking behavior, and reception has spread cross-platform โ€” but HN sentiment splits hard on Aleph Alpha's credibility and the decisive adoption artifacts (quants, llama.cpp support) still do not exist. Heat rises to high on attention grounds, not belief: valve-eligible multi-platform top-decile spread, 96.7th-percentile current rate despite the launch burst cooling (124โ†’30 pts/h), and a still-climbing HN thread โ€” while the adoption hypothesis itself remains unverified with a new credibility headwind.

What is this?

Aleph Alpha has released Kolibri-1, an Apache-2.0 open-weight mixture-of-experts language model with 78B total parameters, 3.46B active per token, and a context window of up to 1M tokens, per its own announcement and Hugging Face model card, with a published technical report. The supplied search results contain no independent coverage of Kolibri-1 itself โ€” nothing corroborating the claimed strong LocalLLaMA reception, nor any quantizations, local deployments, or harness integrations โ€” so the adoption question at the heart of this case is unconfirmed in either direction (a search miss is not evidence of fading). What the results do establish is the competitive context: Kolibri-1's spec profile (small-active-param MoE, 1M context, permissive license) is now a crowded 2026 pattern, matched or exceeded by Qwen3.6-35B-A3B, Qwen3.8-27B, Tencent Hy4 preview, DeepSeek V4 Flash, Step 5 Preview, Poolside Laguna S2.1, and China Telecom Xing4.0, with the Qwen coverage framing the real contest as who 'can run reliably, economically, and compliantly on real hardware' โ€” and citing ecosystem size (300K+ Qwen derivatives) as the moat a new entrant must overcome. The snippets add no background on Aleph Alpha as a company beyond the case's own 'known lab' framing.

Why it matters to Scott

Aleph Alpha has independently shipped the exact combination Scott's canon and practice already occupy โ€” a permissively licensed compact sparse-MoE with a 1M-token window aimed at local/agentic inference โ€” which makes Kolibri-1 a dated, home-testable instance of the thesis his Inference Field ebook argues (what million-token windows are for) and a direct test of his dumb-zone/attention-budget skepticism, with the practical twist his hardware-aware-inference concept governs: the 78B total-parameter weight footprint, not the 3.46B active, decides whether a quantization can ever sit on gamepc's Ollama endpoint. But everything load-bearing is still announcement-class by his own evidence ladder โ€” the grounding confirms no independent evals, quants, or deployments corroborate the claimed reception โ€” and the crowded sibling field (Kimi Linear, Ling-3.1, Naive N0.5, Qwen) means the resolution will turn on ecosystem pull rather than specs, exactly the capability-symmetry extension: if 1M context survives on his rig, dumb-zone's typical-cutoff claim weakens; if it fades like its siblings, the ecosystem-as-moat claim strengthens.
ip:source.the-inference-field-ebookip:concept.dumb-zoneip:concept.attention-budgetip:concept.evidence-class-ladderip:concept.capability-symmetrydev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcradar:kimi-linear-local-validationradar:ling-31-flash-releaseradar:naiveai-n05-flash-releaseradar:quasar-438b-european-modelradar:concept.open-weight-modelsradar:concept.long-context-inferenceradar:concept.sparse-moeradar:concept.quantizationradar:concept.llama-cpp
queries asked of Scott's wikis
  • open-weights strategy model sovereignty
  • sparse MoE active parameters inference economics
  • long-context memory cost local hardware
  • open model adoption signals quants llama.cpp
  • coding agent harness open model tool calling
  • vendor release claims tech report evaluation

Measured heat

now 0 pts/hpeak 392 pts/hcomments 0/hpeers p38momentum: steady3 platformsage 85h
points/hour across evidence ยท reading as of 2026-10-07 11:04:53.997080+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-03 12:26 (minted)โญ origin echo-reconstructedAnnouncing Kolibri-1: 78B parameters, 3.46B active, up to 1M tokens of context, Apache 2.0 โ€” with an X announcement and a published tech rep
Aleph Alpha on github (echo) ยท attributed from reddit.post.1wwl7y6 ยท published time unknown
โ€”
10-03 10:43first on hacker news ยท published ยท lag ?Show HN: Germany's new sovereign AI model Kolibri
tejaskumar__
โ€”
10-03 11:44first on r/LocalLLaMA ยท published ยท lag ?Aleph-Alpha/Kolibri-1 ยท Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0
Nunki08
โ€”
10-05 10:42first on r/singularity ยท published ยท lag ?German lab Aleph Alpha releases Kolibri: a sovereign open-weight model,78B parameters, 3.46B active. Up to 1M tokens of context.
TorturedPoet30
โ€”
10-03 10:43amplified on hacker news ๐Ÿ‘‘hn.story.49943034
tejaskumar__
peak 420 ยท 143 comments ยท 40% of case engagement
10-03 11:44amplified on r/LocalLLaMAreddit.post.1wwl7y6
Nunki08
peak 571 ยท 167 comments ยท 29% of case engagement
10-03 17:22amplified on hacker newshn.story.49946069
yu3zhou4
peak 109 ยท 5 comments ยท 8% of case engagement
10-05 09:16amplified on r/LocalLLaMAreddit.post.1wy3y4t
EveYogaTech
peak 144 ยท 36 comments ยท 7% of case engagement
10-05 10:42amplified on r/singularityreddit.post.1wy5b5g
TorturedPoet30
peak 288 ยท 89 comments ยท 15% of case engagement
10-05 17:15amplified on r/LocalLLaMAreddit.post.1wyeanf
Psychological_Lab955
peak 1 ยท 4 comments ยท 0% of case engagement
3 more amplifiers in ainews.case_chain
10-03 12:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wwl7y6โ€”
10-03 13:38reached heat=high ยท lag ? ยท via ledgerโ€”โ€”

Evidence (10) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditAleph-Alpha/Kolibri-1 ยท Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0
LocalLLaMA
Nunki08571167
๐ŸŸง echo.github โญAnnouncing Kolibri-1: 78B parameters, 3.46B active, up to 1M tokens of context, Apache 2.0 โ€” with an X announcement and a published tech repAleph Alphaโ€”โ€”
๐ŸŸง hnShow HN: Germany's new sovereign AI model Kolibritejaskumar__42012
๐ŸŸง hnKolibri โ€“ Tech Report [pdf]yu3zhou41095
๐ŸŸ  redditYou can now try Aleph Alpha's Kolibri 78B for free online here.
LocalLLaMA
EveYogaTech14436
๐ŸŸ  redditGerman lab Aleph Alpha releases Kolibri: a sovereign open-weight model,78B parameters, 3.46B active. Up to 1M tokens of context.
singularity
TorturedPoet3029290
๐ŸŸ  redditI squeezed Kolibri-1 78B-A3.5B to 20.9 GiB / 2.30 bpw โ€” 59% lower KL than standard IQ2_XS
LocalLLaMA
Psychological_Lab95504
๐ŸŸ  redditLess Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.
LocalLLaMA
kmodi155
๐ŸŸ  redditHosted Kolibri writing Nim code
LocalLLaMA
carlocapocasa82
๐ŸŸ  redditWe gave Aleph Alpha's Kolibri-1 up-to 72 action combinations and put it in Doom. What could go wrong? ๐ŸŽฎ
LocalLLaMA
kmodi02

Interpretation history

Decision trace