K2 Horizon is an open-weight model family released September 3, 2026 by the Institute of Foundation Models (IFM), a frontier lab launched by Abu Dhabi's MBZUAI in May 2025. It spans six Apache 2.0 models β 0.9B, 3.7B, 7B, and 32B dense, a 36B-total/~4B-active sparse model using IFM's MoVA (mixture-of-value attention), and a 375B-A23B sparse model β with day-one GGUF builds, claimed SOTA at the small scales, and an unusually open training lifecycle (data, intermediate checkpoints, training code, and logs slated for release through end of September 2026). IFM also ships 'Uno', a drop-in decoding adapter claimed to give roughly 3Γ lossless speedup (advertised in community posts as up to 5,200 tok/s on the 7B). The supplied coverage stresses that the vendor's benchmark claims are not yet independently corroborated, and the snippets conflict on active-parameter counts (the HF card says ~4B per token; the vLLM recipe says 5.95B active on 37.44B stored).
2026-10-06T22:23:14Z
The substantive_evidence trigger is the llama.cpp support PR (#29535, dense+MoVA) becoming first-class case evidence β but the PR is still unmerged and its thread only re-confirms known findings (prohibitive KV-cache cost; Barni275's quantized-KV rerun puts the 7B below Ornith-1.5-9B/Qwen3.5-9B), so the case's meaning is unchanged: a quiet, corroborated release whose leaning-negative sovereignty test stays dormant on discrete triggers. Low heat holds β the magnitude-valve top-decile spread is the already-priced September launch wave, while current motion is ~2.8 pts/h at 800h age.
2026-10-06T20:42:17Z
evidence attached: reddit.post.1wz7pha β llama.cpp PR adding K2 Horizon dense/MoVA support is concrete tooling-adoption evidence for the model family's local-usability claim.
2026-10-05T05:24:04Z
The velocity-spike triggers dissolve on inspection: the AMA gained two points over four days against a near-zero aged-cohort baseline β a trickle, not a revival β with no new evidence, implementations, outlets, or AMA answers. The case's meaning is unchanged and it now sits definitively past its promised end-of-September artifact window: a quiet, corroborated release whose leaning-negative sovereignty test waits on discrete triggers (llama.cpp/Ollama merge, independent Uno validation, confirmed artifact delivery or slippage, FT confirmation). Polling has been dropped through the quiet ladder; this is adjudication of a noise trigger, not a story.
2026-10-01T21:51:40Z
IFM's first-party AMA keeps the lab visibly engaged, but its captured content is community pressure rather than new answers β KV-cache economics, benchmarks trailing Qwen 3.6, why the training-data open-sourcing is delayed β and the promised end-of-September artifact window has now lapsed without confirmed delivery, sharpening the case's meaning into a leaning-negative live test of sovereign-software-assurance's demonstrated-not-promised bar. The measured uptick (~5 pts/h, 82nd peer percentile) is the AMA as sole mover in an aged cohort β the magnitude-valve spread was the already-priced launch wave β so corroborated state and low heat hold.
2026-10-01T20:35:47Z
evidence attached: reddit.post.1wv8zww β First-party IFM AMA on the K2 Horizon release (funding, data-delay, deployment answers) is direct follow-up material for the open case.
2026-09-29T12:33:23Z
Velocity has flatlined to zero (0 pts/h and 0 comments/h against a ~274/h launch peak, peer percentile 0) and the periphery stopped expanding β no new communities, outlets, or implementations beyond one runtime comment β so the case steps down from accelerating to corroborated: deployable and validated-enough to run, but no longer moving. The only substantive addition is ik_llama.cpp working out of the box (a third inference path that partially bypasses the unmerged llama.cpp PR) plus confirmation that vLLM lacks MoVA and the 7B KV-cache cost is prohibitive; both refine the deployment picture without changing the waiting posture.
2026-09-25T18:49:20Z
grounded: converges/medium β The case's arc independently lands where Scott already argued: vendor small-model SOTA claims are now contradicted by the first independent 16GB benchmark (7B w
2026-09-25T18:40:29Z
The FT mainstream-corroboration lead fails to confirm: its HN thread (hn.story.49847170) is a 7-point, 3-comment stub whose comments never tie 'cheap model taking aim at OpenAI/Anthropic' to K2 Horizon, so it stays an open lead rather than independent press validation. Heat drops to low because the magnitude-valve spread is the already-priced launch wave β current velocity is ~0.83 pts/h against a 274/h peak and this window added no new implementations, communities, or outlets β leaving a waiting posture on discrete triggers (llama.cpp/Ollama merge, an independent Uno speed/quality run, FT confirmation).
2026-09-25T18:28:12Z
evidence attached: hn.story.49847170 β FT mainstream coverage of a cheap model challenging OpenAI/Anthropic appears to be the K2 Horizon story; if confirmed this is independent corroboration of the open accelerating case.
2026-09-23T18:00:37Z
GGUF quantizations across the lineup are now the second implementation line (alongside oMLX) confirming real local deployability, but this look added only engagement drift and repetitive speculation β no new independent speed, quality, or cost validation. The episode has earned its spread (Reddit front-page threads, HN, AA listing, third-party quants and runtime ports), so it graduates to accelerating on implementation breadth rather than post volume, while the load-bearing claims (Uno 5,200 tok/s losslessness, KV-cache economics, agent-task quality) remain open.
2026-09-22T21:23:03Z
evidence attached: reddit.post.1wnky7x β The newly available GGUF quantizations are a concrete deployment artifact and independent corroboration that expands K2 Horizon's local usability.
2026-09-18T01:39:15Z
Uno discussion adds neither an independent speed/quality result nor a verified local execution path; benchmark enthusiasm and benchmaxxing suspicions are equally unsupported here. Cool attention while retaining the diffusion adapter as a concrete evaluation lead, not a demonstrated deployment breakthrough.
2026-09-17T19:22:34Z
The reported K2-Horizon-7B-Uno release adds a concrete diffusion-adapter route to faster inference, rather than another parameter-efficiency claim. Its advertised 5,200 tokens/s and lossless speedup warrant renewed attention, but the supplied coverage establishes neither consumer-hardware gains nor preserved agent-task quality.
2026-09-17T19:21:49Z
evidence attached: reddit.post.1wj2hsm β Independent coverage points to a released K2 Horizon artifact and a potentially consequential 5,200-token-per-second diffusion-augmented inference claim, though the speed claim remains unvalidated.
2026-09-16T11:28:10Z
A commenter adds timeout and task-completion claims to the already-known unfavorable 7B comparison, but is not the benchmark author and provides no independent run or supporting logs. This reinforces caution about agent-task economics without establishing a new implementation result or changing the family's status as a bounded-context evaluation candidate.
2026-09-16T04:23:28Z
The new hybrid-attention comment identifies a possible confound in the Qwen comparison, but supplies neither verified architecture details nor a rerun. It reinforces the need for matched memory and quality measurements without changing K2 Horizon's status as a feasible local evaluation candidate rather than a demonstrated cost breakthrough.
2026-09-15T22:10:22Z
A firsthand benchmark reports K2 Horizon 7B running entirely within 16GB VRAM, countering blanket claims that its attention architecture makes local use impractical, while reporting quality far behind Qwen3.8-27B on that setup. This strengthens the deployability evidence without establishing competitive task economics; the supplied excerpt omits the context length, complete quantization settings and numerical results.
2026-09-15T21:23:10Z
evidence attached: reddit.post.1wha6sv β Independent local benchmark materially contextualizes K2 Horizon's practical quality and fit on 16GB VRAM.
2026-09-15T00:26:35Z
The latest discussion amplifies the already-known long-context memory caveat without supplying measurements or a new deployment result. Unsupported claims of prohibitive serving costs do not establish that bounded-context local use is impractical, leaving the family an evaluation candidate rather than a demonstrated cost breakthrough.
2026-09-14T17:49:55Z
The latest 7B post adds a preliminary coding-use anecdote, not a completed implementation result or independent confirmation of the repeated benchmark claims. It keeps the small models worth bounded-context testing but does not establish suitability for low-VRAM deployment or change the cost thesis.
2026-09-14T17:24:41Z
evidence attached: reddit.post.1wg82rd β A relatively well-engaged independent report and linked GGUF provide useful corroboration of K2 Horizon's small-model local-inference potential, while noting architecture and KV-cache concerns.
2026-09-14T14:30:20Z
A firsthand 7B loading report extends the long-context memory caveat beyond the sparse 36B: llama.cpp projected roughly 89,543 MiB at a reported default 500k context, not measured usage at a typical workload. Reported Artificial Analysis results make the smaller models more interesting to test, but neither parameter-efficiency plots nor extreme-context memory projections establish practical deployment economics.
2026-09-14T14:22:21Z
evidence attached: reddit.post.1wg4a0u β Independent local-inference analysis materially contextualizes K2 Horizon by showing its KV-cache design can dominate effective memory requirements.
2026-09-14T12:36:52Z
The new discussion recycles small-model benchmark enthusiasm and skepticism without adding an implementation result or a credible benchmark contradiction. The runtime-specific evaluation opportunity remains, but this amplification does not strengthen the deployment thesis and warrants cooling attention.
2026-09-14T12:21:57Z
evidence attached: reddit.post.1wg0vqz β User reports strong real-world interest in K2 Horizon's small models, adding adoption context to the open-model release case but not independent validation.
2026-09-13T09:22:15Z
A self-identified oMLX support implementer now reports using K2 Horizon as a daily-driver replacement for gpt-oss-120b, strengthening the case for practical Mac deployment despite problems on other configurations. This makes runtime-specific evaluation more promising, but the claimed speed and tool-calling advantages remain an unmatched, interested firsthand account rather than comparative validation.
2026-09-13T02:31:09Z
A new firsthand report adds a concrete long-context memory constraint to the existing runtime risks: the 36B MoVA reportedly does not fit in 32GB VRAM at 128k context because of KV-cache costs. This further limits the low-cost local-deployment thesis without establishing a universal hardware limit; the same user's praise for the 7B does not demonstrate comparable long-context feasibility.
2026-09-11T21:32:29Z
New firsthand reports turn runtime maturity from a general uncertainty into a concrete deployment risk: one 16GB AMD user reports severe offload-dependent slowdowns, alongside separate loading and output-format problems. This weakens the near-term low-cost deployment thesis without disproving model quality; the truncated RPC account does not establish multi-device performance.
2026-09-11T21:22:00Z
evidence attached: reddit.post.1wdrvd4 β Independent local use reports expose serious MoVA runtime and multi-device performance limitations relevant to the open model's practical deployability.
2026-09-09T23:35:01Z
Refreshed discussion adds no substantive evidence beyond the known release, mixed coding anecdotes, and preliminary MLX throughput report. Practical local execution remains corroborated, but competitive quality, deployment savings, upstream integration, and delivery of the promised training artifacts remain unresolved.
2026-09-07T23:29:28Z
The latest community inventory repeats previously known model sizes rather than establishing a family expansion, and flags the 32B offering as a stage-one checkpoint rather than a finalized model. Practical local execution remains corroborated, but competitive quality, deployment savings, and delivery of the promised training artifacts remain unvalidated.
2026-09-07T23:22:19Z
evidence attached: reddit.post.1wa6tyr β This independently corroborates the reported K2 Horizon release and its expanding family of dense and sparse open models.
2026-09-07T11:27:45Z
The attached university coverage reinforces release provenance, not independent validation of capability or deployment savings; it comes from IFMβs own institution. Practical local execution remains corroborated, while competitive quality, inference economics, and delivery of the promised full training lifecycle remain unsettled.
2026-09-07T11:22:47Z
evidence attached: hn.story.49596565 β First-party MBZUAI coverage independently corroborates the K2 Horizon open-model release.
2026-09-06T22:33:02Z
No new evidence beyond prior repetitive amplification; local execution via MLX remains the only corroborated fact, while competitive quality and sparse-inference economics claims are still unvalidated. The case has plateaued and continues cooling with no fresh signal.
2026-09-04T22:29:37Z
The refreshed discussion adds no reproducible quality results, broader runtime support, or measured memory, context-scaling, and cost evidence. Local execution remains corroborated, while competitive capability and sparse-inference economics remain unsettled and the episode continues to cool.
2026-09-04T17:30:49Z
Refreshed comments only repeat known integration questions, benchmark skepticism, and the earlier coding concern; they add no reproducible quality or inference-economics evidence. Practical local execution remains corroborated, but the broader competitiveness thesis is still unsettled and the episode is cooling.
2026-09-04T13:36:14Z
Refreshed comments add implementation questions and repeat the existing coding concern, but provide no new reproducible quality, memory, prefill, context-scaling, or cost evidence. Practical local execution remains corroborated while competitive quality and sparse-model economics remain unsettled.
2026-09-04T12:31:14Z
Independent MLX throughput measurements, combined with separate hands-on coding reports, move K2 Horizon from launch-only interest to corroborated practical local execution. The result is slower than comparable A3B models and still lacks memory, context-scaling, quality, and cost controls, so the broader competitive-quality and inference-economics thesis remains unsettled.
2026-09-04T12:22:45Z
evidence attached: reddit.post.1w71oyi β Independent community benchmark results materially bear on K2 Horizonβs practical local-inference throughput and validate the open-model deployment case.
2026-09-04T11:28:25Z
The refreshed discussion adds no reproducible capability tests, runtime integration results, or measured memory, throughput, and cost evidence. Existing mixed anecdotes still leave K2 Horizon as a hands-on evaluation candidate rather than a corroborated quality or inference-economics shift.
2026-09-04T10:29:24Z
The refreshed comments remain repetitive amplification of mixed anecdotes and benchmark skepticism, without reproducible capability tests, runtime integration results, or measured inference economics. K2 Horizon remains a relevant local evaluation candidate but has not advanced toward corroboration.
2026-09-04T09:29:01Z
Refreshed comments and engagement add no reproducible capability tests, runtime integration results, or measured memory, throughput, and cost evidence. The mixed anecdotes still make K2 Horizon a local evaluation candidate, not a corroborated quality or inference-economics shift.
2026-09-04T08:27:33Z
The refreshed discussion adds no reproducible evaluations, runtime integration results, or measured memory, throughput, and cost evidence. Existing mixed anecdotes leave K2 Horizon worth testing locally but do not advance the capability or sparse-inference economics thesis.
2026-09-04T07:40:44Z
The refreshed discussion adds no reproducible capability comparisons, working integration results, or measured memory, throughput, and cost evidence. The mixed qualitative reports remain insufficient to validate either competitive quality or the sparse modelβs local-inference economics.
2026-09-04T06:28:17Z
A qualitative hands-on report now says the 36B-A4B model performs decently on agentic coding tasks comparable models can handle, adding a tentative positive signal alongside the earlier negative 3.7B anecdote. Without prompts, measurements, reproducible comparisons, or inference-cost data, this is not enough to corroborate the broader capability and economics hypothesis.
2026-09-04T05:26:55Z
Refreshed discussion and engagement add no reproducible quality tests, deployment measurements, or working integration evidence. The release remains a relevant local evaluation candidate, but its capability and sparse-inference economics remain uncorroborated.
2026-09-04T04:29:12Z
The new thread identifies integration readiness and benchmark-harness comparability as practical blockers, but supplies no measured quality or inference-economics results. K2 Horizon remains a worthwhile local evaluation candidate pending upstream support, usable quants, and reproducible comparisons.
2026-09-04T04:22:09Z
evidence attached: reddit.post.1w6t0a9 β shared external link with case evidence
2026-09-04T03:33:34Z
The refreshed discussion remains repetitive amplification and benchmark skepticism, adding no reproducible quality tests or measured local-inference economics. K2 Horizon remains a relevant evaluation candidate but has not moved toward corroboration.
2026-09-04T02:28:21Z
The refreshed comments add no reproducible benchmarks, deployment measurements, or independent implementation results. K2 Horizon remains a relevant local-testing candidate, but its quality and sparse-inference economics are still uncorroborated.
2026-09-04T01:28:40Z
The refreshed discussion remains repetitive amplification and benchmark skepticism, with no new reproducible evaluation or measured local-inference results. The preliminary coding concern remains anecdotal, so the release is still an evaluation candidate rather than a corroborated capability or economics shift.
2026-09-04T00:28:10Z
The refreshed discussion adds no independent benchmarks, reproducible tests, or deployment measurements beyond the already-known coding anecdote. K2 Horizon remains a hands-on evaluation candidate, but neither its quality nor sparse-inference economics has advanced toward corroboration.
2026-09-03T22:45:40Z
A first independent hands-on report now gives a preliminary negative signal for the 3.7B modelβs coding reliability, replacing the prior absence of any user evaluation. The undocumented single-prompt anecdote is too weak to establish capability or inference economics, but it increases the need for reproducible testing rather than promotion.
2026-09-03T21:38:57Z
Refreshed comments continue to amplify the releaseβs openness and question its benchmark positioning without adding independent quality, throughput, memory, cost, or deployment evidence. K2 Horizon remains a worthwhile local evaluation candidate, but the case has not advanced toward corroboration.
2026-09-03T20:36:46Z
The refreshed comments and engagement remain repetitive discussion of openness and benchmark credibility, not independent evidence of model quality or local-inference economics. The case still merits hands-on evaluation but has not advanced toward corroboration.
2026-09-03T19:44:24Z
The refreshed discussion remains amplification of the releaseβs openness and benchmark skepticism, with no independent quality, throughput, memory, cost, or deployment evidence. The case remains a hands-on evaluation candidate, but this update adds no momentum toward corroboration.
2026-09-03T18:54:00Z
Refreshed discussion still repeats the openness appeal and benchmark skepticism without adding independent evaluations, deployment measurements, or implementation results. The release remains a timely local-testing candidate, but its quality and inference-economics claims are unchanged and uncorroborated.
2026-09-03T16:44:45Z
The added Hacker News discussion broadens awareness but contributes no independent benchmarks, deployment measurements, or implementation results. The release remains a credible hands-on evaluation candidate whose quality and inference-economics claims are uncorroborated.
2026-09-03T16:23:29Z
evidence attached: hn.story.49551760 β shared external link with case evidence
2026-09-03T15:54:35Z
Refreshed comments remain repetitive amplification of the launch, openness claims, and benchmark skepticism rather than independent validation. K2 Horizon is still a relevant hands-on evaluation candidate, but there is no basis to promote it beyond watching.
2026-09-03T14:40:16Z
Refreshed discussion reinforces practical interest in the smaller models and unusually open training lifecycle, but also sharpens skepticism about benchmark competitiveness and benchmaxxing. No independent quality, throughput, memory, or cost evidence has arrived, so the case remains an evaluation candidate rather than a corroborated capability shift.
2026-09-03T14:27:42Z
grounded: converges/medium β IFMβs open checkpoints, reproducibility emphasis, and 4B-active sparse model converge with Scottβs software-sovereignty position and his preference for deployab
2026-09-03T14:25:12Z
case created β The first-party announcement and released Hugging Face artifacts establish a substantive open-model launch with active local-inference interest.