2026-10-11 17:10 UTC

Magic claims its V5 pretraining recipe exceeds leading open-weight base models’ compute efficiency by more than tenfold on its held-out loss evaluations, potentially substantially lowering the compute needed for competitive base-model training.

state: watchingheat: lowuncertainty: highnovelscott: lowpretraining-efficiency ai-infrastructure training-systemsMagicFireworks

What is this?

In a September 8, 2026 research update, the Magic Team claims its pretraining recipe is more than ten times as compute-efficient as those of leading open-weight base models. Magic says it matches DeepSeek V4 Pro Base with roughly 50 times fewer FLOPs (approximately $0.5 million of GB200 compute), and that a larger run costing roughly $4 million outperforms publicly available open base models on perplexity evaluations. These are company-reported comparisons and scaling-law cost estimates, not independently verified results or demonstrated downstream coding-agent performance; the supplied snippets do not establish the V5 designation or Fireworks’ involvement.

Why it matters to Scott

No meaningful intersection with Scott’s specific claims or active builds is established: cheaper pretraining is only adjacent to his Cost of Cognition position, which concerns useful cognitive output, and Magic’s company-reported loss comparisons establish neither cheaper inference nor better agent performance. The radar already tracks related efficiency claims in “looping-20b-token-efficient-pretraining,” but the supplied hits do not show it tracking this Magic development.
radar:concept.training-efficiencyradar:looping-20b-token-efficient-pretrainingradar:concept.model-evaluation
queries asked of Scott's wikis
  • algorithmic efficiency versus compute scale frontier training economics
  • capital barriers model ownership open weights strategy
  • perplexity held-out loss versus downstream agent capability
  • training cost estimates FLOPs hardware utilization reproducibility
  • base model pretraining recipes versus post-training coding performance

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 818h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-07 14:00⭐ origin echo-reconstructedMagic says its recipe is '>10x more compute-efficient' than leading open-weight base models and matches DeepSeek V4 Pro Base using approxima
Magic Team on blog (echo) · attributed from hn.story.49613072
—
09-08 16:59first on hacker news · published · +27.0h>10x More Efficient Pretraining
ronfriedhaber
—
09-08 16:59amplified on hacker news 👑hn.story.49613072
ronfriedhaber
peak 118 · 61 comments · 100% of case engagement
09-10 15:22our radar first saw it · +73.4hdiscovery anchor: hn.story.49613072—
pace: p73 vs 519 stories at the 720h mark (now 818h old) — ahead of microsoft-vibevoice-streaming-asr (1.0x), behind openai-collective-cyber-defense (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn>10x More Efficient Pretraining
Retrieved article excerpt

Open article · Retrieved 2026-09-10T15:27:44.022441+00:00

>10x More Efficient Pretraining Research update on compute-efficient pretraining and scaling to trillion-parameter models. Magic Team , on September 8, 2026 Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. After compounding for … a while …, our pretraining recipe is now >10x more compute-efficient than that of leading open-weight base models. We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. We continued scaling 10x (~$4M) and meaningfully outperformed all publicly available open base models on perplexity evals. By the scaling laws in Figure 1, training a model this capable would cost >$100M under DeepSeek V4 Pro’s recipe (and this is ignoring how much data exists). Of course, we won’t stop scaling there. We believe pretraining, agentic RL, and long-context are sufficient to build superhuman coding agents and automate AI R&D. We started with long-context . Today’s blog post is about pretraining. bits per byte (lower is better) Private Code Repos i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.185 0.204 0.223 0.242 0.261 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.206 bpb DeepSeek V4 Pro: 0.202 bpb Kimi K2: 0.205 bpb Nemotron 3 Ultra: 0.203 bpb V5 e21: 0.246 bpb V5 e22: 0.222 bpb V5 e23: 0.202 bpb V5 e24: 0.194 bpb 29x DSv4 Flash 48x DSv4 Pro 31x Kimi K2 47x Nemotron 3 Ultra Heldout Research Papers i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.36 0.41 0.46 0.52 0.57 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.415 bpb DeepSeek V4 Pro: 0.404 bpb Kimi K2: 0.421 bpb Nemotron 3 Ultra: 0.406 bpb V5 e21: 0.527 bpb V5 e22: 0.463 bpb V5 e23: 0.407 bpb V5 e24: 0.383 bpb 24x DSv4 Flash 45x DSv4 Pro 41x Kimi K2 35x Nemotron 3 Ultra Reasoning on heldout math problems i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.53 0.64 0.76 0.87 0.98 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.710 bpb DeepSeek V4 Pro: 0.678 bpb Kimi K2: 0.695 bpb Nemotron 3 Ultra: 0.619 bpb V5 e21: 0.890 bpb V5 e22: 0.770 bpb V5 e23: 0.647 bpb V5 e24: 0.587 bpb 72x DSv4 Flash 127x DSv4 Pro 58x Kimi K2 15x Nemotron 3 Ultra 6·N·D training FLOPs Figure 1 : Pretraining scaling laws against training compute, comparing to leading available open-weight base models . 1 We measured bits-per-byte loss (a metric that normalizes out differences in tokenizers) on heldout data and fit a scaling law to project how much compute is needed to reach a given level of capability. Better training compute efficiency means stronger models at all budgets. We evaluated the latest available open-weight base models 2 from DeepSeek, Moonshot (Kimi), and NVIDIA. Base models for Claude, Gemini, GPT-n, and many others aren’t openly available, but Kimi K3 and Meta’s Muse Spark indicate a 2.5x and 3.3x gain over Kimi K2, respectively. We evaluated logprobs for open models in both vLLM and SGLang on both GB200 and GB300 and found issues with some backends in the process. For further confirmation, we partnered with Fireworks to verify baseline logprobs in their in-house inference engine. Since models can learn their training parser’s characteristics, we built our eval sets using a different parser/OCR than the one our pretraining pipeline uses. Evaluating generalization To measure generalization, we evaluated loss on heldout data (Figure 1). Our code evals consist of our own codebase and private codebases we acquired from other startups. For reasoning evals, we generated CoT and step-by-step walkthroughs to heldout, private math problems using Kimi K3 and filtered for correct answers. For text and research, we used recent, low-citation research papers. We removed vendored OSS code and any document with a matching 96-character window of normalized text or Jaccard similarity above a sensitive threshold compared to our training data. 3 Evaluating knowledge In addition to generalization, we are interested in testing our model’s knowledge in key domains to identify gaps in our dataset. For example, we can decompose our heldout research text eval set by subject. bits per byte (lower is better) Heldout Computer Science Papers i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.37 0.43 0.48 0.54 0.59 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.450 bpb DeepSeek V4 Pro: 0.437 bpb Kimi K2: 0.458 bpb Nemotron 3 Ultra: 0.434 bpb V5 e21: 0.548 bpb V5 e22: 0.481 bpb V5 e23: 0.425 bpb V5 e24: 0.400 bpb 62x DSv4 Flash 120x DSv4 Pro 108x Kimi K2 69x Nemotron 3 Ultra Heldout Engineering Papers i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.359 0.409 0.458 0.507 0.556 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.413 bpb DeepSeek V4 Pro: 0.403 bpb Kimi K2: 0.421 bpb Nemotron 3 Ultra: 0.405 bpb V5 e21: 0.517 bpb V5 e22: 0.456 bpb V5 e23: 0.405 bpb V5 e24: 0.383 bpb 26x DSv4 Flash 48x DSv4 Pro 49x Kimi K2 38x Nemotron 3 Ultra Heldout Math Papers i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.32 0.37 0.43 0.48 0.54 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.365 bpb DeepSeek V4 Pro: 0.354 bpb Kimi K2: 0.367 bpb Nemotron 3 Ultra: 0.360 bpb V5 e21: 0.492 bpb V5 e22: 0.426 bpb V5 e23: 0.367 bpb V5 e24: 0.342 bpb 12x DSv4 Flash 21x DSv4 Pro 16x Kimi K2 22x Nemotron 3 Ultra Heldout Physics Papers i Lower bits per byte is better. Training compute is in FLOPs on a logarithmic axis. The curve fits our current recipe; the dashed segment is a projection beyond the largest run. 0.38 0.43 0.49 0.54 0.60 10 21 10 22 10 23 10 24 10 25 DeepSeek V4 Flash: 0.433 bpb DeepSeek V4 Pro: 0.422 bpb Kimi K2: 0.440 bpb Nemotron 3 Ultra: 0.425 bpb V5 e21: 0.552 bpb V5 e22: 0.487 bpb V5 e23: 0.431 bpb V5 e24: 0.406 bpb 16x DSv4 Flash 29x DSv4 Pro 29x Kimi K2 24x Nemotron 3 Ultra 6·N·D training FLOPs Figure 2 : Effective-compute per research area . By collecting granular buckets of content (e.g. documentation of a particular software tool or key papers in alignment research) we can get even more precise signals. Unlike for our generalization eval, we don’t want to fully remove much of this information (e.g. key papers in a field) from the pretraining corpus, but we still need to avoid rewarding sequence memorization 4 . To do this, we reworded/summarized these documents using a third-party frontier LLM. To avoid overfitting to granular evals, we created and evaluated them once per model generation; the ones below were made last week. Magic’s goal is to build the best model for coding and autonomous AI R&D. To intentionally balance data mixing trade-offs, we also evaluate domains we deprioritize (e.g. facts about notable people, local news, or sports/events). Our recipe vs. Best open model per eval effective-compute multiplier vs. best open model per eval Eval multipliers are ranked within each panel on a logarithmic axis. Values above 1 favor our current recipe; values below 1 favor the baseline. 0.01x 0.1x 1x 10x 100x SWE & AI R&D SWE AI R&D 0.01x 0.1x 1x 10x 100x Local news, world knowledge & law Local news World knowledge Law Figure 3 : Effective-compute multiplier across domains . We fit scaling laws on eval sets across 167 domains and show compute efficiency gains per dataset. No shortcuts In late 2024, we trained a small dense model with an architecture designed for very long context windows. Our initial pretraining scale-ups kept blowing up in a wide variety of ways. We learned quickly that we had to build a stable foundation first. Smooth convergence, low-precision training quality equivalent to FP32, fast and stable infra, correct hyperparameter scaling rules. And most importantly: hunt the bugs. Once we had that in place, we needed to find enough compute efficiency improvements to close the gap to the frontier with less compute. We had a few big bets to start with, but our progress ended up being the multiplicative result of tens of changes across model architecture, optimizer, training objective, and data curation. NanoGPT speedruns provide a fast feedback cycle to evaluate new ideas, but we found that many things that improve tiny models don’t improve big models. Similarly, we found that some features present in most LLMs can be deleted without harming large scale performance. To evaluate each model, optimizer, or data change, we train 3 models spanning 2 orders of magnitude of compute. We consider a change worth keeping if its power law fit suggests it will help at scale. Every few weeks, we scaled up to 1/10th of our hero scale and every few months we ran a full-scale hero run (V3, V4, V5 in Figure 4). Compute efficiency relative to V2, with competitor reference lines. 1x 10x 100x 1000x V2 Late '24 V3 Early '26 V4 July '26 V5 Sep '26 Compute efficiency vs. our V2 Nemotron 3 · 24x Kimi K2 · 10x DSv4 Pro · 5.6x x1.9 x11 x24 505x vs V2 Figure 4 : Acceleration of our pretraining research progress . To sanity check how pretraining loss translates to post-RL performance, we ran a short math RL run with a 16k CoT budget (Figure 5). All of our RL starts directly from the base model without SFT or distillation. Math pass rate over RL training compute, with competitor reference lines. 0% 25% 50% 75% 100% 0.01% 0.1% 1% 10% 100% RL compute (as % of the model's own pretraining FLOPs) Pass rate GPT-6 Astra · 100% Claude 5.1 · 98% Muse Spark 1.3 · 88% Gemini 3.8 · 83% Kimi K3 · 75% Grok 4.6 · 71% DSv4 Pro · 65% Inkling · 50% Qwen 3.8 · 45% GLM-5.3 · 45% Nemotron 3 · 27% V5 (e24) · 72% V5 (e23) · 31% Figure 5 : Pass@1 on heldout competition math problems during low-compute RL . 5 FLOPs are 6·N·D, as in Figure 1. What’s next Our pretraining and long-context work is now quite mature. We’ll now scale long-horizon RL, training agents to keep learning after deployment through long-context. We’re also putting significant work towards alignment training techniques that present robust theoretical properties. And last but not least, we look forward to releasing the thing! Concrete problems we’re tackling include: Exploration and credit assignment in long-horizon RL (and systems work to scale up). Alignment training against narrowly elicited latent knowledge . 6 Further improvements to pretraining. We are likely the smallest team in the world training trillion parameter models. The impact a single person with strong judgement can have has never been higher. If you want to help build aligned superintelligence, consider joining . Footnotes We report 6·N·D in place of exact training flops, where N is the activated parameter count and D is the pretraining token count each report states. Sequence-dimension (e.g. attention, etc.) cost makes up a minority of the FLOPs for these (and our) models but depends on the exact sequence length distribution used. These aren’t reported for all public models, so we opted for the 6·N·D approximation to avoid guessing. The “6” appears because (add+mul) * (fwd+bwd*2). We derived active parameter counts by downloading the checkpoints’ safetensors headers from HuggingFace, reading each tensor’s shape, and adding up the sizes of all active parameters except token embeddings and MTP heads. 6·N·D for each model: Model N D 6·N·D Current_e24 (ours) - - 1.63e24 Current_e23 (ours) - - 1.58e23 Current_e22 (ours) 
ronfriedhaber11861
🟧 echo.blog ⭐Magic says its recipe is '>10x more compute-efficient' than leading open-weight base models and matches DeepSeek V4 Pro Base using approximaMagic Team——

Interpretation history

Decision trace