2026-10-11 18:01 UTC

open-weight-models

band: hotmomentum: stable score: 1.0
temperature history

Episodes (11)

Chinese regulators will turn consultations with Alibaba, ByteDance, Zhipu, and other firms into restrictions on overseas downloads of leading Chinese open-weight models while preserving access through hosted APIs.
expiredconvergesscott: medium
Vercel claims open-weight models reached 56% of its AI Gateway token volume but only 14% of spending in August 2026, helping lower average token prices by 23.2% and strengthening the economic case for workload-specific model routing.
resolvedconvergesscott: medium
Nathan Lambert's written Congressional testimony claims Chinese open-weight models now dominate open-model usage โ€” roughly twice US Hugging Face downloads (3.2B vs 1.6B) and over 80% of OpenRouter's ~80T weekly open-model tokens while lagging closed frontiers by only 2-5 months โ€” establishing China as the de facto center of the open-weight ecosystem.
resolvedknownscott: medium
Bloomberg reports, citing people familiar with the matter, that Harvey's gross margins swung from about +50% in January to โˆ’50% by June 2026 under frontier-model token costs and turned positive again after it shipped a Kimi K3-based custom model โ€” confirmation would make negative application-layer margins a demonstrated driver of professional-agent vendors' shift to open-weight models.
corroboratedconvergesscott: high
Benchmark author mauricekleine's Nonobench v1.2 reports that among 43 tested LLMs no open-weight model solves the new 20ร—20 Hard mode (0/10) while GPT-6 Astra posts the first perfect 15ร—15 run (30/30) and Opus 5.5 scores 8/10; if the open-weight shutout holds as open-weight releases are added, it stands as a measured open-frontier gap on long-constrained verifiable reasoning, and any open model clearing 20ร—20 refutes it.
watchingconvergesscott: medium
Anthropic's Frontier Red Team claims Zhipu's open-weight GLM-5.3 autonomously builds end-to-end cyber exploits at near-Mythos-Preview level while its safeguards fall to simple bypasses 64โ€“100% of the time (corroborated by NIST CAISI's 'most cyber-capable open-weight model to date' assessment), and whether this disclosure โ€” with browser 0-days already disclosed to maintainers โ€” draws concrete vendor, buyer, or governance responses to open-weight cyber risk resolves the episode.
corroboratedconvergesscott: high
Baseten announces a partnership making open-weight models (GLM-5.3 Flash, Kimi K3) natively usable inside Codex with spend counted against OpenAI commits; sustained enterprise adoption โ€” or OpenAI narrowing the openness โ€” decides whether OpenAI's flagship coding harness has become a multi-model platform that commoditizes model choice.
watchingconvergesscott: high
Mindgard claims jailbreaks of Moonshot's Kimi K2.6 and K3 Swarm produce bioweapon and assassination guidance despite guardrails, and Moonshot โ€” whose internal review began only after BBC contact โ€” either ships a documented fix that establishes open-model bio-uplift as a live cross-border safety issue, or the report joins the pile of unremediated jailbreak disclosures.
seedconvergesscott: high
Per the Axios scoop, Reflection AI โ€” the Nvidia-backed startup founded by ex-DeepMind researchers and led by Misha Laskin โ€” is about to release its first open-weight frontier model positioned as a US answer to DeepSeek and Qwen (backed by $7B+ in committed compute); the weights, license, and benchmarks actually landing โ€” and whether builders adopt it as a competitive US open-weight option for coding agents and regulated procurement โ€” resolve whether a US lab has re-entered the open-weight frontier or this stays an announcement of an announcement.
resolvedconvergesscott: high
vox-deorum's controlled CivBench claims GLM-5.3 now beats Opus 5.5 at long-horizon Civilization V play while Qwen-3.8-27B stays competitive; CivBench becoming a cited reference benchmark for long-horizon strategic planning across frontier and open-weight models โ€” or its GLM-over-Opus ranking failing replication โ€” resolves it.
watchingnovelscott: high
Nace.ai claims its open-source Drex 1.5 9B model tops the JevBench 0.3.1 leaderboard and achieves the lowest latency on OpenRouter among sub-9B decision models, potentially establishing it as the reference routing model for agent orchestration.
seedconvergesscott: high

Trajectory notes