2026-10-11 16:38 UTC

IFM claims its released K2 Horizon family combines competitive model quality, multiple open local-inference sizes, and a sparse 36B model with 4B active parameters, potentially lowering the cost and improving the reproducibility of capable local deployments.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumopen-models local-inference mixture-of-expertsIFM

What is this?

K2 Horizon is an open-weight model family released September 3, 2026 by the Institute of Foundation Models (IFM), a frontier lab launched by Abu Dhabi's MBZUAI in May 2025. It spans six Apache 2.0 models β€” 0.9B, 3.7B, 7B, and 32B dense, a 36B-total/~4B-active sparse model using IFM's MoVA (mixture-of-value attention), and a 375B-A23B sparse model β€” with day-one GGUF builds, claimed SOTA at the small scales, and an unusually open training lifecycle (data, intermediate checkpoints, training code, and logs slated for release through end of September 2026). IFM also ships 'Uno', a drop-in decoding adapter claimed to give roughly 3Γ— lossless speedup (advertised in community posts as up to 5,200 tok/s on the 7B). The supplied coverage stresses that the vendor's benchmark claims are not yet independently corroborated, and the snippets conflict on active-parameter counts (the HF card says ~4B per token; the vLLM recipe says 5.95B active on 37.44B stored).

Why it matters to Scott

The case's arc independently lands where Scott already argued: vendor small-model SOTA claims are now contradicted by the first independent 16GB benchmark (7B well behind Qwen3.8-27B, reported 40-minute agent-task timeouts) β€” a dated receipt for evaluation-driven development β€” while the open weights and promised-but-undelivered training artifacts keep testing sovereign-software-assurance's 'demonstrated, not merely promised' bar. It also bears on his own stack: GGUF quants make bounded-context evaluation on gamepc/Ollama possible today (blocked in Ollama only until the llama.cpp PR merges), and the 36B-A4B's low-bandwidth MoVA results decide whether it earns the cheap end of his model barbell β€” so it stays medium, an evaluation candidate, until the Uno speed claim or upstream merge triggers action.
ip:concept.model-barbellip:concept.evaluation-driven-developmentip:framework.sovereign-software-assuranceip:framework.scout-senior-splitdev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcradar:adaptive-kv-cache-streamingradar:afm3-prompt-conditioned-pruningradar:adaptive-speculative-decoding-300-gpu
queries asked of Scott's wikis
  • model barbell economics small local models vs frontier API
  • open weights software sovereignty local deployment strategy
  • mixture-of-experts sparse active parameters inference cost
  • KV cache long-context memory limits local inference hardware
  • Ollama llama.cpp GGUF quant evaluation harness gamepc stack
  • open training data code reproducibility model replication

Measured heat

now 0 pts/hpeak 42 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 914h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-03 14:25 (minted)⭐ origin echo-reconstructedIntroduces the K2 Horizon model family as a radically open release with frontier-level performance, multiple sizes, and reproducible trainin
IFM on blog (echo) Β· attributed from reddit.post.1w68rj6, reddit.post.1w67wso Β· published time unknown
β€”
09-03 13:47first on r/LocalLLaMA Β· published Β· lag ?IFM/K2-Horizon-MoVA-36B-A4B-GGUF Β· Hugging Face
jacek2023
β€”
09-03 15:36first on hacker news Β· published Β· lag ?K2 Horizon: Frontier Performance, Radically Open
karimf
β€”
09-03 13:47amplified on r/LocalLLaMAreddit.post.1w67wso
jacek2023
peak 242 Β· 102 comments Β· 9% of case engagement
09-03 14:19amplified on r/LocalLLaMAreddit.post.1w68rj6
Few_Painter_5588
peak 604 Β· 191 comments Β· 20% of case engagement
09-03 15:36amplified on hacker news πŸ‘‘hn.story.49551760
karimf
peak 335 Β· 127 comments Β· 21% of case engagement
09-04 03:31amplified on r/LocalLLaMAreddit.post.1w6t0a9
edward-dev
peak 118 Β· 47 comments Β· 4% of case engagement
09-04 11:27amplified on r/LocalLLaMAreddit.post.1w71oyi
DerTomsn
peak 1 Β· 12 comments Β· 0% of case engagement
09-07 10:33amplified on hacker newshn.story.49596565
andy99
peak 3 Β· 0 comments Β· 0% of case engagement
11 more amplifiers in ainews.case_chain
09-03 14:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1w68rj6β€”
pace: p96 vs 519 stories at the 720h mark (now 914h old) β€” ahead of mentria-bonsai27b-webgpu-inference (1.2x), behind meta-muse-personal-agent (1.0x)

Evidence (18) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditIntroducing K2 Horizon: Frontier Performance, Radically Open
LocalLLaMA
Few_Painter_5588602191
🟠 redditIFM/K2-Horizon-MoVA-36B-A4B-GGUF · Hugging Face
LocalLLaMA
jacek2023242102
🟧 echo.blog ⭐Introduces the K2 Horizon model family as a radically open release with frontier-level performance, multiple sizes, and reproducible traininIFMβ€”β€”
🟧 hnK2 Horizon: Frontier Performance, Radically Openkarimf335127
🟠 redditHas anyone already tried IFM's new K2-Horizon-MoVA-36B-A4B?
LocalLLaMA
edward-dev11847
🟠 redditK2-Horizon-MoVA-36B-A4B-MLX-4bit: up to 49.1 tok/s for local inference β€” llm-bench.io
LocalLLaMA
DerTomsn012
🟧 hnK2 Horizonandy9930
🟠 redditIFM/K2-Horizon-*
LocalLLaMA
PhilippeEiffel45
🟠 redditIs anyone using K2-Horizon-MoVA-36B-A4B? If yes, what is the usecase?
LocalLLaMA
DerTomsn1928
🟠 redditThe new k2 horizon models seem like an absolute beast
LocalLLaMA
Eyelbee27384
🟠 redditK2 Horizon lineup is out on AA, and once again AA plots are misleading.
LocalLLaMA
crusaderky7529
🟠 redditFor the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index.
LocalLLaMA
Uncle___Marty63067
🟠 redditI benchmarked IFM/K2-Horizon-7B on 16GB VRAM
LocalLLaMA
Barni2752725
🟠 redditIFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss
LocalLLaMA
Zulfiqaar9418
🟠 redditquants for K2-Horizon are now available
LocalLLaMA
jacek20234819
🟧 hnThe cheap new AI model taking aim at OpenAI and Anthropicmichaelhillaert135
🟠 redditAMA about K2 Horizon, Meet our team from IFM
LocalLLaMA
aya-ifm123137
🟠 redditmodel : add K2 Horizon dense and MoVA support by bitalov · Pull Request #29535 · ggml-org/llama.cpp
LocalLLaMA
jacek2023176

Interpretation history

Decision trace