2026-10-11 17:12 UTC

Independent reruns will determine whether MiniMax M3 Medium reproducibly achieves about 73.17% F1 on DeepSearchQA and approaches leading proprietary models on practical deep-research tasks.

state: expiredheat: lowuncertainty: highconvergesscott: mediumopen-models deep-research model-evaluationMiniMaxYou.com

What is this?

MiniMax M3 is a model from Shanghai-based MiniMax, described in the supplied results as combining coding capability, a one-million-token context window, and multimodal input; sources differ on whether its weights had actually been released at the time of evaluation. The case claims a 73.17% F1 score on DeepSearchQA, a benchmark for systematic web exploration, multi-source synthesis, and search completion, based on an artifact reporting an adjusted score of 0.7316989793753246. The supplied snippets do not independently establish that score, its proximity to GPT-5 High, You.com’s role, or that independent reruns have confirmed the result, so reproducibility remains the event to verify.

Why it matters to Scott

If independently reproduced, the result would support Scott’s capability-symmetry position by showing a potentially open model approaching proprietary performance on practical research-agent work, while directly inviting the trace-backed, vendor-neutral comparison he advocates. It could affect model selection and provide a concrete benchmark fixture, but the supplied evidence does not yet establish the score or reproducibility.
ip:concept.capability-symmetryip:concept.capability-auditdev:concept.trace-backed-agent-comparisonip:concept.evaluation-driven-developmentradar:concept.open-modelsradar:concept.research-agentsradar:concept.agent-benchmarksradar:concept.model-evaluation
queries asked of Scott's wikis
  • deep-research agent evaluation and reproducibility
  • benchmark harnesses for search and multi-source synthesis
  • open-weight models versus proprietary research agents
  • F1 metrics for agentic research quality
  • long-context models in RAG and knowledge workflows
  • practical model evaluation beyond vendor benchmarks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnMiniMax M3 Medium hits 73.17% F1 on DeepSearchQA, near GPT-5 HighEdwardIrby10
🟧 echo.other ⭐The primary result artifact reports generatedAt 2026-08-21 and adjusted averageScore 0.7316989793753246 (73.17%) for minimax/minimax-m3 acroYou.com / Edward Irby——

Interpretation history

Decision trace