open-models
band: hotmomentum: stable
score: 1.0
Episodes (177)
Trajectory notes
- 2026-10-06T14:47:11Z: deepseek-v41-flash-release closed (absorbed) — The expanding runtime and hardware ecosystem converges with Scott’s hardware-aware local-inference work, while the mixed agent results reinforce his position that economics must be measured per successful model-plus-harness task ra
- 2026-10-04T18:28:09Z: minicpm5-2b-release closed (disproved) — As supplied, this is another instance of Scott’s already-held Model Perishability position: model progress warrants replaceable backends and re-evaluation, with his gamepc/Ollama stack providing a concrete testing destination. The radar
- 2026-10-04T07:52:14Z: heretic-model-unrestriction-tooling closed (absorbed) — Scott's own canon already carries the exact claim this case demonstrates — Architecture, Not Vibes and Guardrail Illusion hold that model-level behavioural controls are removable probability barriers and cannot serve as th
- 2026-09-30T11:53:55Z: spark-x25-small-model-release closed (faded) — Scott already distinguishes nominal million-token capacity from usable attention-residence and treats memory pressure, accelerator placement, and runtime support as explicit local-inference policy; those positions are carried by Th
- 2026-09-29T19:24:12Z: deepseek-v4-flash-vision-release closed (absorbed) — If the weights and licence are genuinely available, DeepSeek is extending a consequential model family toward Scott’s existing combination of self-hosted GPU inference, hardware-aware deployment, and vision-equipped agent har
- 2026-09-27T13:33:57Z: glm-53-coding-cyber-validation closed (absorbed) — Independent benchmark lines and Mouse’s task-level, token-accounted run converge with Scott’s Model-Plus-Harness Benchmark Unit and trace-backed comparison doctrine: GLM-5.3 is a credible, economical test candidate, not a valid
- 2026-09-25T04:53:39Z: alibaba-qwen4-announcement closed (superseded) — The alleged Qwen 4 announcement is unsupported and should not be promoted beyond rumor; Scott’s Discussed Is Not Deployed and Evidence Class Ladder already require that evidence ceiling. The substantiated Qwen3.8-27B development
- 2026-09-25T02:27:50Z: yandex-aliceai-foundation-base closed (window-closed) — The release is adjacent to Scott’s gamepc local-serving work, but the supplied evidence establishes neither compatibility with his runtimes nor a performance or resource advantage that would change what he builds; the cust
- 2026-09-24T17:54:06Z: xiaomi-mimo-26-release closed (absorbed) — Independent hands-on evidence now converges with Scott's canon: the 'benchmaxxed' accusation and senior-SWE testing finding benchmarks don't translate to real coding are exactly his Benchmarking the Wrong Unit / Model-Plus-Harness argu
- 2026-09-10T18:00:54Z: applied-compute-training-serving-platform closed (faded) — AC2’s managed checkpoint, dataset and deployment lifecycle repeats the production discipline Scott already holds in Production AI Systems and the 12-Factor Agents Framework; the supplied evidence does not establish inte