Aleph Alpha has released Kolibri-1, an Apache-2.0 open-weight mixture-of-experts language model with 78B total parameters, 3.46B active per token, and a context window of up to 1M tokens, per its own announcement and Hugging Face model card, with a published technical report. The supplied search results contain no independent coverage of Kolibri-1 itself โ nothing corroborating the claimed strong LocalLLaMA reception, nor any quantizations, local deployments, or harness integrations โ so the adoption question at the heart of this case is unconfirmed in either direction (a search miss is not evidence of fading). What the results do establish is the competitive context: Kolibri-1's spec profile (small-active-param MoE, 1M context, permissive license) is now a crowded 2026 pattern, matched or exceeded by Qwen3.6-35B-A3B, Qwen3.8-27B, Tencent Hy4 preview, DeepSeek V4 Flash, Step 5 Preview, Poolside Laguna S2.1, and China Telecom Xing4.0, with the Qwen coverage framing the real contest as who 'can run reliably, economically, and compliantly on real hardware' โ and citing ecosystem size (300K+ Qwen derivatives) as the moat a new entrant must overcome. The snippets add no background on Aleph Alpha as a company beyond the case's own 'known lab' framing.
Aleph Alpha has independently shipped the exact combination Scott's canon and practice already occupy โ a permissively licensed compact sparse-MoE with a 1M-token window aimed at local/agentic inference โ which makes Kolibri-1 a dated, home-testable instance of the thesis his Inference Field ebook argues (what million-token windows are for) and a direct test of his dumb-zone/attention-budget skepticism, with the practical twist his hardware-aware-inference concept governs: the 78B total-parameter weight footprint, not the 3.46B active, decides whether a quantization can ever sit on gamepc's Ollama endpoint. But everything load-bearing is still announcement-class by his own evidence ladder โ the grounding confirms no independent evals, quants, or deployments corroborate the claimed reception โ and the crowded sibling field (Kimi Linear, Ling-3.1, Naive N0.5, Qwen) means the resolution will turn on ecosystem pull rather than specs, exactly the capability-symmetry extension: if 1M context survives on his rig, dumb-zone's typical-cutoff claim weakens; if it fades like its siblings, the ecosystem-as-moat claim strengthens.
ip:source.the-inference-field-ebookip:concept.dumb-zoneip:concept.attention-budgetip:concept.evidence-class-ladderip:concept.capability-symmetrydev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcradar:kimi-linear-local-validationradar:ling-31-flash-releaseradar:naiveai-n05-flash-releaseradar:quasar-438b-european-modelradar:concept.open-weight-modelsradar:concept.long-context-inferenceradar:concept.sparse-moeradar:concept.quantizationradar:concept.llama-cpp
queries asked of Scott's wikis
- open-weights strategy model sovereignty
- sparse MoE active parameters inference economics
- long-context memory cost local hardware
- open model adoption signals quants llama.cpp
- coding agent harness open model tool calling
- vendor release claims tech report evaluation
2026-10-07T01:41:55Z
The pre-registered fade criterion fires: the fourth vendor-adjacent demo (kmodi/Tesseracted Doom, 0 pts / 50% ratio โ even affiliated promotion no longer lands) plus zero leading signal of a 4-bit quant, vLLM/upstream support, or any genuine deployment inside the window, with attention at 0.17 pts/h (0.04% of peak) by hour 85. The velocity-spike and valve-eligible triggers are static-cohort artifacts of the launch burst on a frozen 3-platform set, not periphery expansion, so heat stays low despite the loud spread reading. Meaning shifts from 'open adoption test' to 'resolved non-stick datapoint': the fade lands on the ecosystem-as-moat side of Scott's capability-symmetry claim, and the only live residue is his optional 1M-context home test via the lone 20.9 GiB IQ2_XS quant.
2026-10-06T23:36:36Z
evidence attached: reddit.post.1wzdn1r โ Third-party real-time interactive use of Kolibri-1 at 39ms median is a toy but direct adoption datapoint for the release case.
2026-10-06T12:17:32Z
The 'Kolibri writing Nim code' post attached as adoption evidence is independent in authorship (carlocapocasa's 3code harness) but runs entirely on Tesseracted's affiliated free endpoint โ hosted-usage demo, not deployment โ correcting the attach framing just as the Breakout post was corrected. The adoption ledger is back to zero genuine multi-party uptake: one zero-traction 2-bit quant plus vendor-adjacent demos, with attention at ~1.4% of peak and cooling; the case firms onto the fade trajectory and its days-scale window now closes within ~1โ2 days absent a 4-bit quant, vLLM/upstream llama.cpp support, or any real deployment.
2026-10-06T11:33:25Z
evidence attached: reddit.post.1wyzdjx โ Early third-party hosted usage showing Kolibri writing working Nim code is exactly the adoption evidence the open case tracks.
2026-10-06T04:41:37Z
The velocity flag resolves to a slow-burn hosted-demo thread (~4 pts/h at 65h age), not spread, and the new Breakout 'adoption' evidence is by kmodi/Tesseracted โ the same affiliated party running the free endpoint โ so it downgrades from early community adoption to vendor-adjacent promotion, correcting the prior attach framing. With the lone quant still at zero traction, the case firms along the fade trajectory; resolution now waits only on whether a 4-bit quant or vLLM/upstream llama.cpp support lands inside the remaining days-scale window.
2026-10-06T03:33:52Z
evidence attached: reddit.post.1wynppg โ Third-party open-weight experiment running Kolibri-1 as a 25ms real-time decision loop is early community-adoption evidence the release case is waiting on.
2026-10-05T19:04:10Z
The case's decisive adoption test registered its first hit: a third-party GGUF/llama.cpp quantization now exists (IQ2_XS, 2.30 bpw, 20.9 GiB), converting 'zero quants, no llama.cpp path' from absent to nascent โ though at score-0/33%-ratio traction it is a toe in the water, not a wave. Meaning shifts from 'pure fade wait' to 'fade with one lifeline': the days-scale window narrows but stays open, and resolution now turns on whether a usable 4-bit quant or vLLM/upstream llama.cpp support follows.
2026-10-05T17:32:20Z
evidence attached: reddit.post.1wyeanf โ First third-party quantization of Kolibri-1 with measured KL/top-token gains is exactly the community-quantization adoption evidence the case's resolution criterion tracks.
2026-10-05T11:30:19Z
Periphery trickled rather than expanded: the new second-community post is thin and dismissive ('numbers don't justify the hype', outdated-benchmark skepticism, plus astroturfing/downvote-buying drama), the hosted endpoint ticks at small-post rates, and the measured 'accelerating'/88.5th-percentile reading is again a small-cohort artifact โ the fastest object is a 42-point post running 4.3 pts/h against a 0.67 peer baseline while aggregate flow sits near 4% of the 363-pt peak. Meaning firms along the already-set fade trajectory: conversion still absent at ~49h, so the case now just waits out its days-scale quant/llama.cpp window with quiet โ window-closed as the default end; cooling to low despite valve eligibility because the additions are thin, partly negative, and not implementations.
2026-10-05T11:25:24Z
evidence attached: reddit.post.1wy5b5g โ Kolibri launch spreading to a second community with a notable abstention-training detail โ direct corroboration context for the adoption-tracking case.
2026-10-05T09:34:33Z
Launch attention has spent itself โ current velocity ~2.8 pts/h against a 276 peak, both HN threads flat, and the measured 'accelerating' momentum label is an artifact of one small new post, not real spread. The periphery added exactly one thin item: a third-party free hosted chat endpoint, the case's first external distribution artifact, accompanied by a new credibility ding (Aleph Alpha's own benchmark chart ranks Qwen3.5 above Qwen3.6, contradicting external consensus). Meaning shifts from 'attention peaking while decisive artifacts absent' to 'attention spent, conversion still absent' โ the case now rides a fade trajectory toward non-stick unless quants or llama.cpp/vLLM support land within days; cooling to medium despite valve eligibility because the periphery is static, not expanding.
2026-10-05T09:23:01Z
evidence attached: reddit.post.1wy3y4t โ Third-party free hosted chat endpoint for Kolibri-1 is an early distribution/adoption signal bearing directly on the case's adoption test.
2026-10-03T19:58:37Z
Meaning sharpened from 'will adoption come' to 'attention is peaking while the decisive artifacts are absent': the official tech report is now attached as primary artifact (confirming 78.1B total, EnglishโGerman bilingual sovereign positioning) but adds nothing independent, while engagement accelerated instead of fading โ Reddit ~390/121 and HN Show HN ~406 at a 99th-percentile current rate. Yet the discussion content stays benchmark-comparison class, with community consensus forming around 'Qwen 3.5 35B-tier / worse than Qwen 3.6 35B at twice the size', an insider's self-interested coding/agentic praise, and still zero quants, llama.cpp/vLLM support, or harness integrations. State stays watching โ one independent line (the fp8 hands-on) cannot pass corroborated on engagement alone; heat stays high because the conversion window for the decisive artifacts is open now.
2026-10-03T19:26:18Z
evidence attached: hn.story.49946069 โ Aleph Alpha's official tech report is the primary artifact substantiating the open Kolibri-1 release claims.
2026-10-03T13:38:04Z
The episode crossed from announcement-only to first independent signal: a hands-on local fp8 run (~170 tok/s on a 96GB-class RTX Pro 6000) confirms hardware feasibility and speed while flagging verbose/overthinking behavior, and reception has spread cross-platform โ but HN sentiment splits hard on Aleph Alpha's credibility and the decisive adoption artifacts (quants, llama.cpp support) still do not exist. Heat rises to high on attention grounds, not belief: valve-eligible multi-platform top-decile spread, 96.7th-percentile current rate despite the launch burst cooling (124โ30 pts/h), and a still-climbing HN thread โ while the adoption hypothesis itself remains unverified with a new credibility headwind.
2026-10-03T13:24:17Z
evidence attached: hn.story.49943034 โ HN front-page Show HN coverage of the Kolibri release โ same episode, early spread signal for the adoption watch.
2026-10-03T12:35:51Z
grounded: converges/medium โ Aleph Alpha has independently shipped the exact combination Scott's canon and practice already occupy โ a permissively licensed compact sparse-MoE with a 1M-tok
2026-10-03T12:26:12Z
case created โ A first-party open-weight release from a known lab with unusual specs (3.46B active, 1M context, Apache 2.0) and strong day-one LocalLLaMA reception is a bounded episode whose adoption is resolvable.