StepFun, a previously non-frontier Chinese AI lab, released Step 5 Preview around September 18–20, 2026. Artificial Analysis lists it on its Intelligence Index at 44 — the same score AA's own Kimi K3 (max) page shows — at roughly one-third Kimi's token price ($1.00/$2.70 per 1M input/output vs. K3's $3.00/$15.00), placing it on AA's cost–capability Pareto frontier. HN discussion quoting the official announcement describes a sparse MoE (600B total / 27B active) with 1M context and vision input, and StepFun has shipped Step Code, an MIT-licensed first-party coding-agent CLI built around the model. Caveats the snippets themselves raise: several other sources report Kimi K3 at 57 on the AA index (likely a different index version or configuration), Step 5 Preview's weight-access and license are listed as 'Pending' by third-party trackers after an apparently accidental early fork of BF16 weights (now 404, official weights promised October 15), and a demo video was credibly critiqued as reusing an existing project — so the capability-parity claim rests on one evaluator's snapshot, not independent replication.
StepFun independently demonstrates Scott's model-perishability/capability-symmetry thesis with dated receipts — a new Chinese lab reaching Kimi K3's measured tier at one-third price — and its MIT Step Code CLI continues the Chinese-lab first-party-open-harness lineage (Z.ai, MiniMax, DeepSeek) already on the radar. It bears on what he'd build: a direct candidate for his LiteLLM cost tiers if parity survives independent evaluation, a local-inference candidate when the promised Oct 15 weights land (27B-active MoE vs gamepc's constraints), and the unexplained 44-vs-57 AA index discrepancy is live material for his version-bound-assessment position.
ip:concept.model-perishabilityip:concept.capability-symmetrydev:concept.cost-tiered-llm-routingdev:technology.litellmdev:concept.hardware-aware-local-inferencedev:concept.version-bound-ai-assessmentradar:concept.inference-economicsradar:concept.open-weight-modelsradar:concept.agent-harnessesradar:concept.mixture-of-expertsradar:kimi-k3-single-consumer-gpuradar:zcode-open-runtime-release
queries asked of Scott's wikis
- model routing tiers cost per task intelligence threshold
- open-weights Chinese labs frontier parity sovereignty
- first-party coding agent CLI pattern model-owned harness
- sparse MoE local inference hardware requirements deployment
- benchmark index versioning evaluator snapshot reliability
- Pareto frontier price collapses model perishability
now 0 pts/hpeak 17 pts/hcomments 0/hpeers p25momentum: steady3 platformsage 578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-10-09T11:58:53Z
OpenRouter deployment confirms API availability at the quoted price tier; measured heat shows accelerating cross-platform attention (81st percentile) as the Oct 15 open-weights milestone approaches. The price–quality equivalence claim still rests on a single evaluator snapshot (AA 44 vs reported 57 for Kimi) with no independent benchmark replication, and the demo critique remains unaddressed.
2026-10-09T04:52:44Z
evidence attached: reddit.post.1x19lvs — shared external link with case evidence
2026-10-09T02:26:53Z
Step 5 Preview is now deployed on OpenRouter, confirming API availability at the quoted price tier; measured heat shows accelerating cross-platform attention (94th percentile) as the Oct 15 open-weights milestone approaches. The price–quality equivalence claim still rests on a single evaluator snapshot (AA 44 vs reported 57 for Kimi) with no independent benchmark replication, and the demo critique remains unaddressed.
2026-10-08T23:06:43Z
evidence attached: hn.story.50007764 — Step 5 Preview appearing on OpenRouter provides deployment evidence for the price-quality hypothesis.
2026-09-24T01:39:24Z
grounded: converges/medium — StepFun independently demonstrates Scott's model-perishability/capability-symmetry thesis with dated receipts — a new Chinese lab reaching Kimi K3's measured ti
2026-09-24T01:35:57Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-23T15:28:49Z
evidence attached: hn.story.49817387 — StepFun's first-party MIT-licensed Step Code CLI productizes Step 5 Preview as a coding agent, materially contextualising the model's practical relevance in the price/quality case.
2026-09-21T23:25:26Z
Cross-platform attention now warrants high heat, but announcement quotations and discussion of purported weights still do not constitute independent validation of the price–quality claim. Reported architecture details sharpen the deployment questions, while the demo critique cautions against inferring coding-agent performance without establishing a benchmark contradiction.
2026-09-20T12:26:09Z
The new fork link introduces a provenance caveat: its poster claims the weights were accidentally released early and that official weights arrive October 15. This is an access lead, not confirmation of usable weights or independent validation of the price–quality claim; modest cross-platform expansion keeps the case worth watching.
2026-09-20T12:21:56Z
evidence attached: reddit.post.1wleycb — The fork is an additional usable artifact confirming early Step 5 Preview weight availability, though it is not the official release.
2026-09-20T05:22:41Z
The latest Hacker News submission adds an announcement lead, but its supplied record contains only a title—not the official announcement or independent evaluation claimed by the attachment rationale. Visibility is broadening modestly without resolving the price–quality comparison or establishing usable weights.
2026-09-20T05:21:48Z
evidence attached: hn.story.49772532 — The official Step 5 Preview announcement directly corroborates the open case about its measured quality and price frontier.
2026-09-20T04:25:41Z
A new Reddit title points to a purported Step-5-Preview-BF16 Hugging Face repository, opening a possible weights-access angle, but the repository itself has not been inspected. This is a discovery lead rather than direct release verification or independent support for the claimed price–quality parity.
2026-09-20T04:21:26Z
evidence attached: reddit.post.1wl6gro — The Hugging Face artifact provides direct release evidence for the already-open Step 5 Preview episode.
2026-09-19T06:22:23Z
The claim has reached Hacker News, but that submission points back to Artificial Analysis rather than providing independent validation. This broadens its visibility without establishing comparable pricing or resolving the benchmark-score discrepancy.
2026-09-19T06:21:47Z
evidence attached: hn.story.49763660 — Artificial Analysis coverage provides independent market-position evidence for Step 5 Preview's claimed quality and price tradeoff.
2026-09-19T05:27:02Z
grounded: converges/low — The reported capability–price parity would align with Scott’s Model Perishability position and offers a potential candidate for his task-aware model routing, ra
2026-09-19T05:23:52Z
origin walked (codex/luna, conf 0.93): anchor reddit.post.1wkd8hx -> echo.other.7e0f749805 by Artificial Analysis
2026-09-19T05:21:49Z
case created — A named model release with a specific reported benchmark and price comparison establishes a bounded episode, though the underlying results and pricing remain unverified here.