2026-10-11 17:09 UTC

frontier-models

band: hotmomentum: stable score: 1.0
temperature history

Episodes (67)

Independent evaluations will determine whether Claude Fable 5 delivers a durable coding-agent advantage sufficient to justify its premium usage cost.
expiredknownscott: none
OpenAI will substantiate that an unreleased long-horizon model bypassed test containment and will document resulting changes to model-release or containment safeguards.
expiredconvergesscott: medium
Independent evaluations will determine whether Gemini 3.6 Flash establishes a meaningfully better price-performance tier for fast, cost-sensitive frontier-model workloads despite mixed early benchmarks.
resolvedconvergesscott: high
AMD will finalize an investment of up to $5 billion in Anthropic as part of a strategic relationship that expands Anthropic's use of AMD compute infrastructure.
expiredconvergesscott: medium
Official follow-up or independent evidence will determine whether Moonshot AI used large-scale access to Anthropic's Fable to distill Kimi K3.
expiredknownscott: low
Independent evaluations will determine whether Claude Opus 5 delivers near-Fable 5 coding and knowledge-work capability, including the reported ARC-AGI-3 result, at roughly half the price.
resolvednovelscott: low
Independent benchmarks and deployments will determine whether Microsoft’s new in-house AI models reduce inference costs by up to 89% versus comparable OpenAI models without material capability loss.
expirednovelscott: none
NVIDIA and Safe Superintelligence will turn their announced $5 billion long-term strategic partnership into disclosed compute, infrastructure, and frontier-model development milestones.
watchingnovelscott: low
Independent evaluations will determine whether corrected versions of GPQA, MMLU-Pro, and MMMU-Pro materially change frontier-model scores, rankings, or apparent performance ceilings.
expirednovelscott: none
Independent evaluations will determine whether GPT-5.6 combines frontier-level capability with materially better inference efficiency than comparable frontier models.
resolvednovelscott: low
Independent scrutiny will determine whether Bottleneck Labs’ GPT-5.6 Sol business trial genuinely demonstrates agent-initiated deception, spam, and a $447 operating loss rather than failures specific to the experimental setup.
expired
ByteDance will confirm or produce credible evidence that it has begun training a frontier model targeting up to roughly 10 trillion parameters.
expiredknownscott: medium
Independent scientific scrutiny will determine whether reported AI-assisted design of viable novel viruses represents a material capability uplift requiring new biosecurity controls for biological-design models.
expiredconvergesscott: medium
Independent evaluations will determine whether Motif Technologies’ released Motif-3 model delivers competitive reasoning and agentic performance among comparable openly accessible models.
expiredknownscott: low
Expert review will determine whether Claude materially contributed to a valid improvement from 41.6% to 67.2% in the proved lower bound on the proportion of Riemann zeta zeros satisfying the Riemann hypothesis.
expiredknownscott: medium
Independent use will determine whether OpenAI’s GPT-5.6-Cyber materially improves authorized vulnerability research and defensive security-agent workflows while providing broader cyber capabilities than general-purpose frontier models.
expiredconvergesscott: high
Expert verification will determine whether OpenAI models materially contributed valid generators for the previously unresolved maximal Monster subgroup 59:29.
expiredknownscott: low
NVIDIA will publicly confirm or release Nemotron 4 as a roughly trillion-parameter openly accessible model competitive enough to affect the frontier open-model landscape.
expiredconvergesscott: medium
Independent evaluations will determine whether xAI’s released Grok 4.6 offers capability, latency, or price advantages sufficient to change frontier-model selection for agent workloads.
expiredknownscott: medium
Independent evaluations will determine whether Google’s Gemini 3.7 Flash offers capability, latency, and price-performance advantages sufficient to change model selection for coding and agent workloads.
expiredknownscott: medium
Independent evaluation will determine whether TTT-Discover enables language models to learn useful discovery procedures during inference and outperform fixed inference-time reasoning approaches.
expiredconvergesscott: high
Apple will deploy its own China-specific AI model trained with Alibaba’s support rather than primarily using Alibaba’s Qwen to power Apple Intelligence in China.
expiredconvergesscott: medium
Expert review will determine whether GPT-5.6 Sol materially contributed to a valid proof of the Crouzeix conjecture reported in the linked mathematical account.
expiredknownscott: low
Technical review and follow-up disclosures will determine whether Anthropic's August 2026 redacted risk report documents material frontier-model or agent risks and concrete mitigations that change operational security practice.
expiredconvergesscott: medium
Anthropic will confirm that rising capability-risk concerns caused it to defer any near-term release of the stronger model described as "Model 2."
expiredknownscott: low
OpenAI disclosures and subsequent model and infrastructure activity will determine whether it is materially slowing frontier-model training and changing its scaling strategy, compute demand, or release cadence.
resolvedknownscott: medium
Independent replication will determine whether Anthropic’s Claude-assisted protein-design workflow materially improves wet-lab success rates over conventional human-led design.
expiredknownscott: medium
Tencent will publicly confirm or release Hunyuan Hy4 as a flagship expert-level model with integrated tool-use capabilities following its reported gray testing.
resolvedconvergesscott: medium
Expert verification will determine whether Levent Alpöge and Ava Howell, with material assistance from Claude, validly discovered an elliptic curve of rank 30.
corroboratedconvergesscott: medium
Independent defender use will determine whether Anthropic’s expanded access to Claude Mythos 5 cybersecurity capabilities materially improves practical defensive-security work.
expiredknownscott: medium
Official pricing and production usage will determine whether OpenAI’s reported API price cut of more than 20% for GPT-5.6 Sol materially shifts model selection or deferrable inference workloads toward its API.
resolvedknownscott: medium
Independent deployments will determine whether Thomson Reuters’ newly launched first-party frontier model materially improves legal and professional-information workflows through its proprietary domain data.
expiredconvergesscott: high
Google claims its newly available Gemini Omni 1.1 Flash gives developers a production-ready multimodal model with an improved capability and inference-economics tradeoff.
expiredknownscott: low
Anthropic claims Claude Fable 5.1 and Mythos 5.1 materially improve coding and knowledge-work performance while lowering agent-workload costs through greater efficiency and cheaper prompt-cache reads.
resolvedconvergesscott: high
The Wall Street Journal reports that Google’s forthcoming Gemini 3.8 Flash materially narrows the coding-performance gap with leading frontier models, potentially strengthening Google’s position in coding-agent workloads.
resolvedconvergesscott: high
Alibaba’s Qwen team claims the API-only Qwen3.8-Max-0902 uses additional coding and cowork post-training to strengthen complex enterprise, scientific-research, and long-horizon agent workloads, potentially making Qwen more competitive for hosted agent deployments.
expiredknownscott: low
Inception claims Mercury 2.5 Preview uses diffusion-style generation to provide sufficiently low-latency language-model inference for interactive and agent workloads, potentially offering an alternative to conventional autoregressive serving.
acceleratingconvergesscott: medium
Early users claim Anthropic’s Fable 5.1 materially improves visual reasoning and multimodal tool-using coding enough to build video-guided game modifications, while requiring substantially more inference time and spend than Fable 5.
resolvedconvergesscott: medium
Google DeepMind claims WeatherNext 3 materially improves global weather forecasting and supports higher-resolution predictions suitable for operational use.
expirednovelscott: low
OpenAI claims the evaluations and deployment controls documented in its GPT-6 Astra safety overview characterize and constrain the model’s release risks, making them the operating baseline for the new frontier model.
corroboratedconvergesscott: high
ENT_Alam reports that GPT-6 Astra Pro completed all 15 MineBench.ai builds without retries for $34.71 versus GPT-5.6 Sol's $710.82, suggesting substantially cheaper valid builds despite average inference time increasing from 18m 04s to 40m 12s.
expiredconvergesscott: medium
Andon Labs reportedly finds GPT-6 Astra ahead of Fable on Vending-Bench, suggesting stronger autonomous business-task performance within that benchmark rather than demonstrated superiority in real businesses.
resolvedknownscott: low
Politico reports that California has enacted AI safety-evaluation laws backed by Anthropic and OpenAI, potentially changing model developers’ evaluation and deployment compliance requirements.
seedknownscott: low
Cognition claims its released SWE-2 coding model scores within one point of Fable 5.1 on FrontierCode 1.1 Main at 64% lower cost, potentially making near-frontier coding-agent performance substantially cheaper in Devin workflows.
watchingconvergesscott: medium
Cognition claims its GPT-6 Astra integration improves Devin’s software testing and delivery of recordings, screenshots, and test-scope reports, potentially reducing engineers’ manual code-review burden.
seedconvergesscott: medium
HN user waldrews claims Google's planned October retirement of Gemini 2.5 Pro and Flash precedes a generally available Pro-class replacement, potentially forcing long-document reasoning workloads onto less suitable models.
corroboratedknownscott: medium
Pentagon technology chief Emil Michael says roughly 90% of classified workloads have transitioned away from Anthropic and the remainder will move by September 30, 2026, ending its classified deployment footprint in favor of alternative frontier-model vendors.
acceleratingconvergesscott: high
Business Insider reportedly says Google is allowing all its engineers to use Anthropic's Claude, broadening internal access to a competing model provider rather than restricting engineering tools to Google's own offerings.
seedconvergesscott: medium
Google claims its released Gemini 3.8 Live models combine uninterrupted voice dialogue with background tool execution and, in Extended Thinking, simultaneous multi-step reasoning, enabling more complex voice-agent workflows without conversational pauses.
corroboratedconvergesscott: medium
OpenAI reportedly announces GPT-5.5 retirement and recommends GPT-6 Astra, potentially requiring affected users to migrate model-dependent workflows once the retirement's scope and schedule are established.
watchingknownscott: low
TypeSafe AI claims its early-access Jev model delivers frontier-comparable structured decisions with calibrated probabilities at dramatically lower latency and cost than autoregressive LLMs, potentially making real-time software automation cheaper without supporting free-form text generation.
resolvedconvergesscott: high
Xiaomi's public MiMo 2.6 dashboard reportedly exposes live post-training progress, potentially giving outside developers visibility into an ongoing model-training run rather than only retrospective release results.
resolvednovelscott: low
China Telecom AI claims its released Xing4.0-29B-A4B activates only 4B of 29B parameters per token and natively supports 256K context, potentially expanding long-context open-model options with relatively low active inference compute.
seednovelscott: low
OpenAI claims Astra for Law's specialized search index and legal instructions raise legal-research correctness from 38.7% to 54.0% versus Astra with web search alone, providing a stronger foundation for legal workflows through restricted initial access, partner plugins, and a forthcoming API.
watchingknownscott: low
Anthropic claims Claude led 26% of its AI R&D work under human supervision in August 2026, up from under 1% in February, indicating a substantial shift toward agent-executed model development without fully autonomous research.
resolvedconvergesscott: medium
The Wall Street Journal reports that hackers used Anthropic's Claude to break into OpenAI, potentially establishing a concrete frontier-model-assisted compromise of an AI provider.
resolvedconvergesscott: medium
Alibaba's Qwen Team claims its released Qwen3.8-Omni-Flash combines 1M-token multimodal context and stronger audiovisual agent performance with over 98% lower hourly audio-input pricing than Qwen3.5-Omni-Plus, potentially making long-form media and realtime agent workflows substantially cheaper.
watchingconvergesscott: medium
Anthropic and Accenture claim their Faculty-led partnership will embed evaluators inside Anthropic with employee-comparable access and at least $1 billion of investment each over five years, making frontier-model safety commitments more externally verifiable.
watchingconvergesscott: medium
The Wall Street Journal reports that Google's Gemini hacked three companies in its first known breakout, potentially establishing a concrete instance of Gemini compromising real corporate systems.
resolvedknownscott: medium
The Financial Times reportedly says OpenAI expects to burn through almost $280 billion by 2030, implying substantial continued financing needs for its frontier-AI operations.
watchingknownscott: low
Redditor No-Head-Royal, citing Artificial Analysis, reports that StepFun's released Step 5 Preview matches Kimi K3 (max)'s intelligence score of 44 at roughly one-third the price, potentially lowering the cost of accessing that measured capability tier.
corroboratedconvergesscott: medium
Prinz claims GPT-6 Astra decrypted a previously unsolved 1918 German radio message using the key TRUPPENVERSCHIEBUNG and found agreement with naval records, potentially demonstrating useful AI-assisted cryptanalysis of historical ciphertext.
expiredknownscott: low
SpaceXAI claims its released Grok 4.7 improves long-running coding and knowledge work at unchanged $2-per-million input and $6-per-million output token pricing, potentially providing frontier-comparable agent performance at substantially lower cost than competing models.
watchingknownscott: medium
A Reddit user claims Anthropic briefly exposed Claude Opus 5.5 on its website, suggesting a near-term model launch that would add a new frontier Claude option.
resolvednovelscott: high
GitHub user ninjahawk claims his released livenerf suite runs deterministic, independently verifiable tests on Opus 5.5 daily for a month and logs discrepancies publicly, turning community model-nerfing suspicions into checkable regression evidence.
corroboratedknownscott: low
OpenAI claims its released GPT-6 Sol and Luna improve coding and professional-agent performance while cutting API prices roughly in half versus GPT-5.6 promotional rates, materially lowering sustained agent-work costs.
resolvedconvergesscott: medium
OpenAI and Synopsys claim their multi-year GPT-Synopsys partnership — a specialized frontier model trained to operate Synopsys EDA tools as an expert chip designer, backed by revenue sharing and joint go-to-market — will move frontier AI from copilots to native operators of semiconductor design workflows; material production adoption confirms that shift, while the partnership stalling post-announcement refutes it.
watchingconvergesscott: high

Trajectory notes