2026-10-11 17:13 UTC

Independent scrutiny will confirm whether Kimi K3 consistently matches or exceeds leading closed models across spreadsheet, web-development, and science evaluations.

state: resolvedheat: lowuncertainty: lowknownscott: mediumkimi-k3 model-evaluation open-modelsMoonshot AIKimiAfterQuery

What is this?

Kimi K3 is a new open-weight model from Beijing-based Moonshot AI, described in the supplied snippets as a 2.8-trillion-parameter mixture-of-experts system with native vision and a context window of up to one million tokens. Moonshot reports frontier-level results, while cited independent evaluations from Arena.ai and Vals AI suggest it is competitive with flagship proprietary models and leads on some spreadsheet, automation, browsing, and web-interface coding benchmarks. The available material does not establish the stronger claim that K3 consistently matches or exceeds the leading closed models: multiple snippets say it still trails Claude Fable 5 and GPT-5.6 Sol overall or in general use, despite setting records on selected benchmarks.

Why it matters to Scott

The radar already tracks Kimi K3 across three episodes on `radar:concept.kimi-k3`. Independent cross-domain scrutiny bears directly on Scott’s Capability Audit and Evaluation-Driven Development positions and could inform his provider benchmark and task-aware model routing, but the supplied evidence does not yet establish consistent superiority or deployable local economics.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:project.remote-execdev:concept.task-aware-model-routingradar:concept.kimi-k3radar:concept.benchmark-integrityradar:concept.open-models
queries asked of Scott's wikis
  • open-weight models closing the frontier gap
  • benchmark leadership versus real-world model utility
  • independent evaluation of coding and agent models
  • open-weight model sovereignty and local inference
  • model distillation and capability transfer
  • multimodal long-context models for coding agents

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (50) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
LocalLLaMA
Charuru606121
🟠 redditKimi K3 is top of nextjs eval
LocalLLaMA
Charuru109796
🟠 redditKimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
LocalLLaMA
Qwen30bEnjoyer31937
🟠 redditKimi K3 landed third on the Intelligence Index, ahead of Opus 4.8, and even GPT-5.6 Sol couldn't take #1 from Fable 5. Weights supposedly drop July 27.
artificial
hero886456043
🟠 reddit@mweinbach (Max Weinbach) recreates macOS 27 with real Liquid Glass and native apps in a web browser with Kimi K3
singularity
AdmirableSelection8130248
🟠 redditKimi K3 tops Frontend Code Arena
singularity
MagicZhang1216244
🟠 redditChinese open-weight model beats Opus 4.8 on some benchmarks, first time this has happened
artificial
roll0ver3213
🟠 redditDoes Kimi K3 change the distillation debate?
singularity
TraditionalHome8852348159
🟧 hnShow HN: Kimi K3 spent nearly 8 hours building this 78-card tarot sitelilyucb20
🟠 redditKimi K3 Benchmarks
LocalLLaMA
WhyLifeIs41277383
🟠 redditKimi K3 weights to be released on the 27th.
LocalLLaMA
Different_Fix_2217413105
🟠 redditKimi K3 beats Opus on paper and sits one point behind Fable. I ran it on real agent tasks- here's real scoresheet.
better_claw
ShabzSparq19115
🟠 redditKIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!
LocalLLaMA
Gohab20011991350
🟧 hnTested Kimi K3 for Codingspeckx266
🟠 redditAccording to Agent Arena Kimi K3 ranks at same level as opus thinking
LocalLLaMA
Terminator8575319
🟠 redditI gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed
LocalLLaMA
DeliciousGorilla12228
🟠 redditMy experience with Kimi K3 after a day of API testing
ClaudeAI
Tarandjpop596
🟠 redditKimi-K3 isn’t quite better than Fable yet, but it’s definitely getting closer.
LocalLLaMA
ImaginaryRea1ity918220
🟠 redditPapers for Stable LatentMoE and Gated MLA?
LocalLLaMA
Ok_Warning214690
🟧 hnClaude Fable 5 vs. Kimi K3: Same results, one-third the cost, 4x slowerBrajeshwar20
🟧 hnKimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTApiotrgrabowski874425
🟧 hnKimi K3: second only to Fable 5 on AA-Briefcasewertyk638
🟧 hnMicrosoft Considers Replacing ChatGPT and Claude with Kimi K3mihau30
🟧 hnCan Chinese AI (Kimi K3) file American tax returns (TaxCalcBench)?michaelrbock136
🟧 hnKimi K3 Is Close to Claude on a Real Coding Taskvincent_s10
🟠 redditKimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI / CAISI.
singularity
socoolandawesome340122
🟠 redditWe gave Fable 5, GPT‑5.6 Sol and Kimi K3 the same six newsroom jobs. Fable edited. Sol complied. Kimi extracted.
ClaudeAI
local___host030
🟧 hnKimi K3 found 19 0days in latest Redis 8.8.0 in 1.5hrsdoanbactam101
🟠 redditA real-world Kimi K3 test: rebuilding a 36-second launch film as editable code
singularity
Tight-Switch819523
🟠 redditKimi K3 across an 8-pass production test: a 75-second code-driven deep-sea film
singularity
gowri1609400
🟧 hnKimi K3 built a Windows XP in browserboveyking5632
🟠 redditKimi K3 gets open weighted tomorrow!
LocalLLaMA
Hot_Example_445649171
🟠 redditKimi K3 countdown has been released
LocalLLaMA
Unusual_Guidance2095536176
🟠 redditOpus 5 High Comes Close, but Kimi K3 Still Leads on Frontend
ClaudeAI
AmbitiousSeaweed10144378
🟧 hnKimi-K3 Releases on HuggingFace 7/27nateb20221375542
🟠 redditK3 Weights are out
LocalLLaMA
HatEducational99656814
🟠 redditKIMI K3’s WEIGHTS ARE OUT!
LocalLLaMA
BritishDudeGuy52196
🟠 redditHere it is boys, The Kimi K3 2.8T
LocalLLaMA
Altruistic_Heat_953118947
🟠 redditK3 Dropped!
LocalLLaMA
jimmystar889673
🟧 hnKimi K3 Tech Reportnekofneko20
🟠 redditKimi-K3 is published on HuggingFace
artificial
BankApprehensive7612355
🟧 hnKimi-K3 Technical Report [pdf]vinhnx358161
🟠 redditKimi K3: Open Frontier Intelligence
singularity
yogthos13713
🟧 hnYou Could Have Come Up with Kimi Delta AttentionAnhTho_FR263108
🟧 hnKimi K3 Architecture Overview and NotesModelForge49999
🟠 redditKimi K3 (Max) takes #1 in the new Code Arena Fullstack rankings, over GPT-5.6 Sol (#2) and Claude Fable 5 (#3)
singularity
KickLassChewGum17619
🟠 redditBenchmarking Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash on Baba Is You
ClaudeAI
pmigdal18028
🟧 hnKimi K3-256kmonneyboi459133
🟠 redditClaude topped business benchmark by lying to suppliers
ClaudeAI
Due-Cup957411
🟧 hnKimi-code not performing wellairbreather21

Interpretation history

Decision trace