Independent reruns will determine whether MiniMax M3 Medium reproducibly achieves about 73.17% F1 on DeepSearchQA and approaches leading proprietary models on practical deep-research tasks.
state: expiredheat: lowuncertainty: highconvergesscott: mediumopen-models deep-research model-evaluationMiniMaxYou.com
What is this?
MiniMax M3 is a model from Shanghai-based MiniMax, described in the supplied results as combining coding capability, a one-million-token context window, and multimodal input; sources differ on whether its weights had actually been released at the time of evaluation. The case claims a 73.17% F1 score on DeepSearchQA, a benchmark for systematic web exploration, multi-source synthesis, and search completion, based on an artifact reporting an adjusted score of 0.7316989793753246. The supplied snippets do not independently establish that score, its proximity to GPT-5 High, You.com’s role, or that independent reruns have confirmed the result, so reproducibility remains the event to verify.
Why it matters to Scott
If independently reproduced, the result would support Scott’s capability-symmetry position by showing a potentially open model approaching proprietary performance on practical research-agent work, while directly inviting the trace-backed, vendor-neutral comparison he advocates. It could affect model selection and provide a concrete benchmark fixture, but the supplied evidence does not yet establish the score or reproducibility.
ip:concept.capability-symmetryip:concept.capability-auditdev:concept.trace-backed-agent-comparisonip:concept.evaluation-driven-developmentradar:concept.open-modelsradar:concept.research-agentsradar:concept.agent-benchmarksradar:concept.model-evaluation
queries asked of Scott's wikis
- deep-research agent evaluation and reproducibility
- benchmark harnesses for search and multi-source synthesis
- open-weight models versus proprietary research agents
- F1 metrics for agentic research quality
- long-context models in RAG and knowledge workflows
- practical model evaluation beyond vendor benchmarks
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-25T00:27:49Z
The result remains an isolated source artifact with no independent rerun or methodological corroboration after the monitoring horizon; the capability-parity implication has not advanced.
2026-08-22T23:37:05Z
No independent rerun or substantive corroboration has appeared; the source artifact remains a testable reported result rather than evidence of reproducible capability parity.
2026-08-22T23:33:32Z
grounded: converges/medium — If independently reproduced, the result would support Scott’s capability-symmetry position by showing a potentially open model approaching proprietary performan
2026-08-22T23:30:37Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49404524 -> echo.other.221b60b94d by You.com / Edward Irby
2026-08-22T23:29:19Z
case created — A published evaluation dataset makes the performance claim testable, but it currently lacks independent validation or broader evidence.
Decision trace
- 08-25 10:27expireThe result remains an isolated source artifact with no independent rerun or methodological corroboration after the monitoring horizon; the capability-parity implication has not advanced.
- 08-25 10:27alert_silentNo new consequential evidence appeared; unchanged engagement and elapsed time do not justify another alert.
- 08-25 10:27alert_routeNo new consequential evidence appeared; unchanged engagement and elapsed time do not justify another alert.
- 08-23 09:37repriceNo independent rerun or substantive corroboration has appeared; the source artifact remains a testable reported result rather than evidence of reproducible capability parity.
- 08-23 09:37alert_silentThis look adds no consequential delta beyond the already-routed artifact publication, so it can wait for an independent reproduction or methodological audit.
- 08-23 09:37alert_routeThis look adds no consequential delta beyond the already-routed artifact publication, so it can wait for an independent reproduction or methodological audit.
- 08-23 09:34alert_shadowA newly published source-of-record artifact reports 73.17% adjusted F1 across 2,688 adjusted trials, nearly matching the separately reported 73.24 score for GPT-5 High. The publication of the result i
- 08-23 09:34alert_routeA newly published source-of-record artifact reports 73.17% adjusted F1 across 2,688 adjusted trials, nearly matching the separately reported 73.24 score for GPT-5 High. The publication of the result i
- 08-23 09:33groundIf independently reproduced, the result would support Scott’s capability-symmetry position by showing a potentially open model approaching proprietary performance on practical research-agent work, whi
- 08-23 09:30promote_anchororigin walk conf 0.98
- 08-23 09:29createA published evaluation dataset makes the performance claim testable, but it currently lacks independent validation or broader evidence.