2026-10-11 18:01 UTC

The Wall Street Journal reports that Google’s forthcoming Gemini 3.8 Flash materially narrows the coding-performance gap with leading frontier models, potentially strengthening Google’s position in coding-agent workloads.

state: resolvedheat: lowuncertainty: mediumconvergesscott: highfrontier-models coding-agents inference-economicsGoogleThe Wall Street Journal
Surfaced 2026-09-02T15:46:41Z — priced heat=high at reprice: The new benchmark post indicates early public experimentation and possible speed gains, but the supplied evidence contains no benchmark values, methodology, access details, or trace-backed results, so it does not independently validate the reported Opus-level coding claim. Imminent release remains the consequential next step.

What is this?

Google is reportedly internally testing a forthcoming model called “Gemini 3.8 Flash Preview” on its Jetski coding platform; Business Insider cites internal images and one employee who found it noticeably better than 3.7 Flash, while Google declined to comment. Google’s released 3.7 Flash is positioned as a lower-cost model for coding and agent workflows, with reported gains in debugging, issue resolution, tool use, and code-generation benchmarks. The supplied snippets do not include the claimed Wall Street Journal report or substantiate comparisons with Opus 5, so the extent to which 3.8 closes the frontier coding gap—and the internal name “Skimaki”—remains unverified here.

Why it matters to Scott

If independently validated, a low-cost Flash model approaching frontier coding performance would reinforce Scott’s model-barbell, task-aware routing, and model-perishability positions while creating a concrete candidate for his LiteLLM-backed Ask workflow. The performance gap is currently unsubstantiated, and the radar already has an open Gemini 3.7 Flash selection case, so this matters chiefly as a potential routing change requiring trace-backed evaluation rather than as an established breakthrough.
ip:concept.model-barbellip:concept.model-perishabilitydev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisondev:project.askdev:technology.litellmradar:gemini-37-flash-model-selectionradar:concept.model-routingradar:concept.inference-economicsradar:concept.coding-models
queries asked of Scott's wikis
  • coding-agent model selection and routing
  • cost-adjusted coding model performance
  • Flash models for high-volume agent workloads
  • token efficiency versus agent task completion
  • frontier model benchmark reliability for coding agents
  • multi-model coding harness portability

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditGoogle back soon? 3.8 Flash competitive with Opus 5 says WSJ
singularity
Charuru18569
🟧 echo.other ⭐The Wall Street Journal’s primary report says Google is preparing Gemini 3.8 Flash, internally called “Skimaki,” with significantly improvedErin Woo——
🟠 redditGemini 3.8 Flash Benchmarks
singularity
Able-Line2683793226
🟠 redditIntroducing Gemini 3.8 Flash and 3.8 Flash Cyber
singularity
Jame9238595
🟠 redditGemini 3.8 flash benchmark in Arfticial analysis
singularity
Expensive_Syrup_652910020
🟠 redditGood, cheap, token hungry
singularity
NoFaithlessness9517232
🟧 hnGemini 3.8 Flashbratao1131638
🟧 hnGemini 3.8 Flash and 3.8 Flash Cybersimonsarris760

Interpretation history

Decision trace