OpenAI launched GPT-6.1 Sol at DevDay on September 29, 2026 — a week after the short-lived GPT-6 Sol — positioning it as a near-Astra model for agentic coding, computer use, and professional work at roughly one-fifth Astra's token price ($2/$10 per Mtok input/output, $0.10/Mtok cached input). The model rolled out same-day in ChatGPT Work, Codex, and the API alongside a new $500/month ChatGPT Pro tier. OpenAI's own benchmarks claim Sol matches Astra on DeepSWE v1.1 at ~1/5 cost, reaches 71.4% on OSWorld 2.0 (vs Astra's 73.5%) at ~1/7 cost per task, and improves on AutomationBench. Independent third-party measurements (PokeBench, Artificial Analysis, MathArena, CafeBench, deterministic photo-to-Blender loops, FrontierMath Tier 4) largely corroborate near-Astra capability at dramatically lower API-side cost on harness-verifiable tasks. However, the same controlled tests contradict the subscription value pitch: Sol 6.1 burns ~38% more $20-plan quota than Opus 5.5 at equal intelligence, and hands-on reports show capability breaks on unmeasured/creative work (Blender game assets, detailed planning refusals) plus two anecdotes of agent-loop turn-count inflation that could invert the token-price advantage exactly on the agent workloads the hypothesis targets. Adoption at scale and independent coding-agent-harness cost/capability results remain unobserved.
| source | object | author | score | comments |
| 🟧 openai ⭐ | DevDay 2026 RecapRetrieved article excerptOpen article · Retrieved 2026-10-10T21:31:18.902689+00:00 September 29, 2026
[Company](https://openai.com/news/company-announcements/)[Product](https://openai.com/news/product-releases/)
# DevDay 2026 Recap
At DevDay 2026, we’re giving people more ways to take on ambitious work and tools to build what comes next.
Share
DevDay 2026 is our biggest yet, with more than 20 major announcements across ChatGPT, Codex, our models, and entirely new forms of working with AI.
We believe AI can help bring about a new renaissance of creativity and discovery. It should give people more time for what matters to them, more freedom to pursue their ideas, and the ability to do things they didn’t think were possible.
Today, we introduced agents that can take on ongoing responsibilities and new ways for people and AI to work together. We also expanded our commitment to an open ecosystem by opening up ChatGPT as a shared surface where humans and agents can collaborate and where developers can directly launch new native experiences to our collective 1.2B weekly users.
Here’s everything we announced.
[New ways of working](https://openai.com/index/devday-2026-recap/#a-whole-new-way-to-work-with-ai)[Build with Codex & the API](https://openai.com/index/devday-2026-recap/#new-tools-for-developers-in-codex-api)[Customize ChatGPT with Plugins](https://openai.com/index/devday-2026-recap/#more-customization-in-chatgpt-with-plugins)[People & AI working together](https://openai.com/index/devday-2026-recap/#people-and-ai-working-together)[Do more with your ChatGPT subscription](https://openai.com/index/devday-2026-recap/#do-more-with-your-chatgpt-subscription)
## A whole new way to work with AI
GPT-6.1 Sol title on a dark, starry background.
Astra Ultrafast title over a starfield, with availability in ChatGPT, Codex, and the API.
Screenshot of a project-specific policy table showing Zero Data Retention with Private Safety Processing and validated external storage.
### Dots
Dots are remarkably capable, always-on agents built to handle everything.They’re a whole new way to work with AI—one that gets to know what matters to you, is always working on your behalf, and takes important work off your plate so you get more of your time and attention back.
Available on Pro and Business Premium in eligible markets. Enterprise, Edu, and Healthcare users can try the beta when their workspace admin enables it; it is off by default. [Learn more](https://openai.com/index/introducing-dots/)
### GPT-6.1 Sol
We’re introducing GPT‑6.1 Sol, a major upgrade to GPT‑6 Sol with exceptionally strong performance on agentic coding, along with computer use, and professional work. It delivers near-Astra intelligence to everyone at a fifth of its standard input and output token prices, giving developers more flexibility to execute complex workflows within the same budget.
Available to all API, Plus, Pro, Business, Enterprise, and Edu users. [Learn more](https://openai.com/index/introducing-gpt-6-1-sol/)
### Ultrafast
Ultrafast is our premium speed tier for workloads where speed matters most. Ultrafast offers up to 8× faster token generation (300 tokens per second) in Codex and up to 6x in the API.
GPT‑6 Astra Ultrafast is available today in the [API(opens in a new window)](https://developers.openai.com/api/docs/guides/ultrafast-mode) and in [ChatGPT Work and Codex(opens in a new window)](https://learn.chatgpt.com/docs/agent-configuration/speed) on Pro 500 and Enterprise plans. GPT‑6.1 Sol Ultrafast is coming soon.
### Private Intelligence
OpenAI Private Intelligence helps businesses use frontier AI with greater confidence that their data is protected. [Zero Data Retention with Private Safety Processing(opens in a new window)](https://developers.openai.com/api/docs/guides/private-safety-processing) enables automated safety reviews without giving OpenAI personnel access to the underlying content. Our preview of Private Inference, coming this fall, combines confidential computing with strict, verifiable controls.
[Contact us](https://openai.com/form/private-intelligence-interest/) about Private Intelligence
## New tools for developers in Codex & the API
Codex cloud projects interface with a purple cloud icon.
Codex CLI terminal interface.
Codex code review interface with a discussion of a code change.
Codex Security Cloud findings dashboard.
Decisions API text on space background.
Diagram of Agents API.
OpenAI and AWS logos on a blue and purple gradient.
### Codex in the cloud
Now developers can run Codex wherever they need it: on a computer, remotely from a phone, or in the cloud from any device. Reusable development environments help tasks start quickly and give your team a shared setup with approved settings and permissions.
Available on Plus, Pro, Business, Healthcare, Education, and Enterprise. [Learn more(opens in a new window)](https://learn.chatgpt.com/docs/cloud)
### A refreshed Codex CLI
The Codex CLI now lets you start and steer tasks with your voice. The new /agents view makes it easier to delegate work and track multiple tasks at once. We’ve also improved everyday workflows to help you edit prompts, resume sessions, and use worktrees, while a cleaner terminal UI makes longer sessions easier to read.
Available to all plans. [Learn more(opens in a new window)](https://learn.chatgpt.com/docs/codex/cli)
### Code Review
The new code review experience in the ChatGPT desktop app makes it easier to review changes across projects. You can read summaries, explore diffs, and ask Codex about potential issues before sharing feedback on GitHub pull requests or GitLab merge requests. With automatic reviews, Codex can also take a first pass in the cloud while you’re away.
Available on all plans. [Learn more(opens in a new window)](https://learn.chatgpt.com/docs/code-review?surface=app)
### Codex Security Cloud
Codex Security Cloud gives defenders a better set of tools to harden their infrastructure. Scan entire GitHub repositories on demand or on a schedule, with ongoing checks of new commits. Codex investigates findings, removes duplicates and prepares fixes in the cloud, even with your laptop closed. It includes access to models offered through [Daybreak Blue](https://openai.com/daybreak/) without a separate Daybreak application.
Available to all Pro, Business, Enterprise and Edu users on desktop and web. [Learn more(opens in a new window)](https://learn.chatgpt.com/docs/security/setup)
### Decisions API
Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent’s next action.
Available in limited preview today with a broad release planned in the coming days.
### Agents API with Computer use
The [Agents API](https://openai.com/index/introducing-the-agents-api/) now supports computer use so developers can build agents that interact with software to complete tasks. It also brings Codex’s multi-agent capabilities, tool search, tool calling, and context compaction into your application. OpenAI runs the underlying infrastructure so your team can focus on building the application.
Available through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise. [Learn more(opens in a new window)](https://developers.openai.com/api/docs/guides/agents-api/tools/computer-use)
### Bedrock Managed Agents, powered by OpenAI
OpenAI worked with Amazon on Bedrock Managed Agents to take the core capabilities of the Agents API and add customization to work natively in AWS and integrate with AWS resources. Now you can use Bedrock Managed Agents to build OpenAI agents that run entirely in AWS. [Learn more(opens in a new window)](https://aws.amazon.com/bedrock/managed-agents-openai/)
## More customization in ChatGPT with plugins
Canva plugin extension shown in the ChatGPT sidebar.
Figma, Shopify, and plugin cards on a blue background.
Three example websites created with Sites.
Diagram linking MCP Events with automations.
### Plugin extensions
We’re opening the platform we use to build ChatGPT features so developers can create their own experiences within ChatGPT. Plugin extensions let you give your plugin a home in the sidebar and build interactive panels where people can work alongside the conversation. You can also create viewers for the file types your product supports.
Available to all plans. [Learn more(opens in a new window)](https://developers.openai.com/plugins/build/extensions)
### Improved plugin creation, submission, and discovery
We’re making it easier to build plugins and help people discover them. Plugin Creator helps you build your plugin while a redesigned submission flow provides clearer feedback. Improved ranking and recommendations help people find relevant plugins in the directory and in conversations. Users choose which plugins to use and approve the access each one receives.
Available to all plans. [Learn more(opens in a new window)](https://developers.openai.com/plugins)
### Sites can now host plugins
You can now add supported ChatGPT plugins to the Sites you build. Teammates in your workspace can use the same app with their own connected data and permissions. We’re also making automations easier to add and manage so Sites can keep shared information up to date.
Available to Business, Enterprise, Healthcare, and Edu plans. [Learn more(opens in a new window)](http://chatgpt.com/features/sites/)
### MCP events for plugin automations
We’re adding support for the [proposed MCP Events specification(opens in a new window)](https://modelcontextprotocol.io/community/working-groups/triggers-events), so plugins can start automations when something happens in a connected app. For example, you can ask ChatGPT to watch for new tasks on a project board. When one comes in, ChatGPT can read the linked documents and draft a plan, even while you’re away.
Available to all plans. [Learn more(opens in a new window)](https://developers.openai.com/plugins/build/mcp-events)
## Improving how people and AI work together
Array of ChatGPT Space windows in a starry field.
Collaborative ChatGPT Page for Q4 event planning with edits from teammates.
Collaborative presentation editor with slide previews and comments.
Schedule a team task interface with team selection.
Microsoft Teams and Slack icons above an @ChatGPT message.
Meeting recap with an @ChatGPT action request.
Shareable profile for Bailey with activity and showcased creations.
### ChatGPT Space
A new home for your team to collaborate with AI to get work done. Create a dedicated space where teammates, ChatGPT, and your dot can build on shared knowledge. ChatGPT can keep your space organized based on instructions you provide, so you can quickly find what you need and pick up where you left off.
Available to all Pro, Business, and Enterprise plans on the ChatGPT desktop app and web, with finding, reading, and sharing pages also available on mobile; mobile creation and editing are coming soon. [Learn more(opens in a new window)](https://chatgpt.com/features/space/)
### Pages
Pages are a new type of document, built for human and agent collaboration. Anything you can do in ChatGPT you can do on a page: write, research, generate charts, create images, or visualize information. Create a page in a conversation, then invite your team to contribute ideas and feedback.
Available to all Pro, Business, and Enterprise plans. [Learn more(opens in a new window)](https://chatgpt.com/features/space/)
### Collaborative slides
Soon you’ll be able to create interactive slides with ChatGPT and your team, from a conversation or your own template. Multiple teammates and agents can edit the deck at the same time and leave comments. Present in ChatGPT or export to PowerPoint or Google Slides with formatting intact.
Available to all Pro, Business, and | OpenAI | — | — |
| 🟠 reddit | OpenAI launches GPT‑6.1 Sol and a $500/month ChatGPT Pro plan OpenAI | ebytes111 | 1 | 3 |
| 🟠 reddit | First impression about gpt6.1 sol OpenAI | Individual_Art_5163 | 2 | 1 |
| 🟠 reddit | Sol 6.1 Max and Blender for games — that just won't cut it. OpenAI | Comfortable-Cat-9611 | 77 | 81 |
| 🟠 reddit | GPT-6.1 Sol beat Pokemon Red's first gym in 244 turns, a new record on PokeBench, for $2.60 OpenAI | VibeCodyH | 157 | 18 |
| 🟧 hn | Artifical Analysis - GPT-6.1 Sol Replaces 6 Sol After 7 Days | Fe2O3 | 2 | 0 |
| 🟧 hn | GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence | theanonymousone | 80 | 99 |
| 🟠 reddit | How are we all finding 6.1 so far? OpenAI | Chemical-Agency-3997 | 12 | 14 |
| 🟠 reddit | 6.1 Sol - So far, pretty impressed OpenAI | therealjerseytom | 175 | 51 |
| 🟠 reddit | What was GPT-6-SOL? OpenAI | Nortixon | 38 | 19 |
| 🟠 reddit | Watched the DevDay stuff yesterday and for the first time I'm actually rethinking my $200 Claude Max sub. ClaudeAI | OccasionNo4703 | 0 | 11 |
| 🟠 reddit | GPT-6.1 Sol is now #1 on MathArena: 86.3% accuracy for $0.94, beating Astra’s 81.9% at $2.26 singularity | 141_1337 | 213 | 17 |
| 🟠 reddit | GPT 6.1 Sol vs GPT 6 Sol vs Opus 5.5 – Testing quota usage (Part 2) ClaudeAI | Background-Web-6312 | 1 | 1 |
| 🟧 hn | CafeBench: Sol 6.1 is great value, but not quite Opus level | zodwick | 5 | 0 |
| 🟠 reddit | GPT 6 Sol lasted 7 days before OpenAI replaced it with GPT 6.1 Sol OpenAI | Top_Knee9687 | 1 | 0 |
| 🟠 reddit | sol 6.1 is actually pretty decent OpenAI | Imapatato12 | 108 | 41 |
| 🟠 reddit | As a Plus subscriber, Sol 6.1 is a game changer OpenAI | penisbike69 | 175 | 59 |
| 🟠 reddit | Photo-to-Blender benchmark: GPT-6 Astra won every photo, GPT-6.1 Sol scored 61 for 36 cents OpenAI | smith2008 | 3 | 7 |
| 🟠 reddit | GPT-6.1 sol (max) scores 100% in frontier math 4 singularity | Southern-Break5505 | 586 | 109 |
| 🟠 reddit | Falling Stars – I didn't miss the code-x ClaudeAI | Peleias | 0 | 3 |
| 🟠 reddit | Downgrading plans as long as I'm using 6.1 Sol? OpenAI | DesiGrit | 13 | 11 |
| 🟠 reddit | gpt 6.1 is the most annoying model i ever used. OpenAI | Simple-Law5883 | 0 | 13 |
| 🟠 reddit | 5.6 Sol is lazy AF singularity | FlamaVadim | 0 | 6 |
| 🟠 reddit | Insane cost difference between Claude Opus 5.5 and Codex Sol 6.1 OpenAI | Electronic_Animal_55 | 52 | 23 |
| 🟠 reddit | Microsoft has accidentally revealed GPT-6.1 Sol uses the same base weights as GPT-6 Sol, but with only 2 inference passes instead of 3. singularity | ResultBackground2450 | 703 | 108 |
| 🟠 reddit | 6.1-sol just disappeared from my codex OpenAI | infinonsen | 2 | 4 |
| 🟠 reddit | Tested OpenAI's "~50% faster" Codex update on my frozen benchmark. Didn't see it (yet) ClaudeAI | Sufficient-Storage87 | 1 | 1 |
| 🟠 reddit | 6x-Sol is so bad but Astra is too expensive OpenAI | Late_Change5029 | 0 | 33 |
| 🟠 reddit | OpenAI rolls out GPT-6 with Intelligent UI — ChatGPT answers can now include interactive tools artificial | winer666 | 23 | 10 |
| 🟧 openai | Asana cuts model costs 76x in browser tests with GPT-6.1 SolRetrieved article excerptOpen article · Retrieved 2026-10-09T19:36:30.956011+00:00 October 9, 2026
# Asana cuts model costs 76x in browser tests with GPT‑6.1 Sol
Using GPT‑6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
[Contact sales](https://openai.com/contact-sales/)
White Asana logo over a blue layered-paper texture.
Company size: Enterprise
Region: North America
Industry: Technology
Products: Codex
76x
Lower estimated model costs with the optimized GPT-6.1 Sol workflow
5x
Faster browser runs with the optimized GPT-6.1 Sol workflow
$0.47
Average estimated model cost with the optimized GPT-6.1 Sol workflow
Loading…
Share
With GPT‑6 Astra in Codex running experiments, Asana optimized its browser agent’s workflow on GPT‑6.1 Sol to run 76x cheaper and 5x faster.
Asana helps customers automate work across business applications through [StackAI(opens in a new window)](https://www.stackai.com/), a platform it [acquired(opens in a new window)](https://asana.com/press/releases/pr/asana-acquires-stackai-adding-cross-system-execution-for-human-agent-teams/e7c73b97-ae8c-4e51-b927-189ccb184146). Using StackAI, customers can build workflows that navigate websites, fill out forms, and gather information without writing code. At Asana’s scale, small inefficiencies in these workflows add up.
Asana’s StackAI CTO, Frank Hidalgo, PhD, set out to make the browser agent faster and cheaper to run. He directed GPT‑6 Astra in Codex to investigate the agent, test improvements, and compare the results. Work he estimates would have taken one to two months by hand took about a week.
[Asana’s 144-run study(opens in a new window)](https://asana.com/inside-asana/cut-browsers-agent-cost) tested GPT‑6.1 Sol and three other frontier models, called here Models A, B and C. The optimized workflow that emerged on GPT‑6.1 Sol averaged $0.47 in estimated model costs and about four minutes per run, 76x cheaper and 5x faster than the original production setup on Model B.
> “This is what teams of humans and agents look like in practice. An engineer set the direction, GPT-6 Astra ran the experiments, and the results went through Command to production. This demonstrates how Asana brings human and agent teams to life.”
—Arnab Bose, CPO at Asana
## Identifying browser-agent inefficiencies with GPT‑6 Astra
To move quickly, Hidalgo started by using GPT‑6 Astra in Codex to map the codebase and explain how the agent built each model request. GPT‑6 Astra discovered that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered, so every request resent that history at full price.
The agent also dropped older screenshots and trimmed text at nearly every step. Each edit altered the history, so caching the history alone would not have helped, and losing those facts could require the agent to revisit pages it had already read.
## From an estimated two months of research to one week with GPT‑6 Astra
Hidalgo reviewed GPT‑6 Astra’s proposed fixes and selected three to test:
- Extending caching to the agent’s browsing history
- Increasing the amount of text it could retain
- Removing screenshots in batches rather than at every step
GPT‑6 Astra began with quick tests to establish which variables mattered. Because the code was not designed for controlled experiments, it then refactored the code so one frontend and backend could support many workflows in parallel, each with its own settings.
Astra conducted the full study: history budgets of 120,000 and 480,000 characters and six caching and screenshot policies, each tested three times on each of the four models (see the table below). The best-performing policy allowed screenshots to accumulate to 20 before cutting back to the most recent one. This kept earlier history unchanged for longer stretches between removals. Combined with the larger history budget, it became the optimized workflow. Each configuration performed the same task: collecting six fields for each of 32 books from a public demo catalog, representative of what some Asana customers run in StackAI.
| **Model** | **Description** | **Price** |
| --- | --- | --- |
| **Model A** | A smaller, less expensive model from another frontier lab, released Fall 2025 | Half the price of GPT‑6.1 Sol |
| **Model B** | The model originally used in production, from the same lab as Model A, released Summer 2026 | Same price as GPT‑6.1 Sol |
| **Model C** | An updated version of Model B, released Fall 2026 | Same price as GPT‑6.1 Sol |
| **GPT‑6.1 Sol** | OpenAI’s model | |
GPT‑6 Astra ran the workflows and examined the requests, usage records and outputs, and separate model sessions reviewed the work. Every session’s requests, data traces and results were recorded in [Command(opens in a new window)](https://asana.com/product/command), Asana’s software delivery platform, so the team could review the complete study afterward. From Command, the findings were turned into tickets, then pull requests, and the changes went to production.
> “This would have taken me one to two months by hand. With GPT-6 Astra in Codex, it took about a week: I’d set a /goal before going to bed and review the results in the morning.”
—Frank Hidalgo, PhD, StackAI CTO at Asana
## Bringing model cost below $0.50 per run
For Model B, the optimization reduced estimated model cost from at least $36.21 (some original runs reached the step limit before they finished) to $1.24 per run, a reduction of 29x. The optimized workflow on GPT‑6.1 Sol was 2.6x cheaper still, at $0.47. Every run in the optimized workflow completed the task and returned the correct answer.
Means of 3 runs. ≥: the baseline includes capped runs, so its mean is a lower bound.
The two right folds compare against Model B optimized. Model B ran in phase 1, Model C and Sol 6.1 in phase 2 of the same study (dotted line).
On GPT‑6.1 Sol alone, with the larger history budget, the new caching and screenshot policy reduced cost 4x, from $1.97 to $0.47 per run. Each call was about 3x cheaper, because 89% of the input came from cache at 5% of the uncached price. Runs also became faster: at least 22.5 minutes on the original setup on Model B, roughly four minutes with the optimized workflow on GPT‑6.1 Sol.
Mean of 3 runs, SD whiskers. ≥: mean includes a capped or unfinished run, so the true value is at least this large.
Bars use the blue theme. Read caching effects against the 480k larger-budget bar.
Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.
Mean of 3 runs, SD whiskers. ≥: mean includes a capped or unfinished run, so the true value is at least this large.
Bars use the blue theme. Read caching effects against the 480k larger-budget bar.
Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.
The investigation also showed how history management affected whether the agent produced an answer at all. Giving GPT‑6.1 Sol more room to retain its browsing history increased the number of runs that produced an answer from three of 18 with the smaller history budget to all 18 with the larger budget, each with the correct answer. For Hidalgo, the business value is giving customers access to faster, more capable models while keeping operating costs sustainable.
> “Cost used to limit which models we could offer customers for these workloads. By making the agent more efficient, we can give customers a better, faster model while lowering our operating costs.”
—Frank Hidalgo, PhD, StackAI CTO at Asana
## Scaling experimentation and product testing
Asana has released the changes to browser navigation in StackAI and is developing tools to make similar experiments easier to repeat. Over time, the team plans to incorporate this testing into the platform’s evaluations, so customers and internal teams can compare cost, runtime, and answer quality when configuring their agents.
> “Shipping speed is no longer the bottleneck; human attention is. We’re close to a world where every engineer is a PM leading a fleet of agents.”
—Frank Hidalgo, PhD, StackAI CTO at Asana
Asana is now using GPT‑6 Astra in Codex to test product features before release: Astra navigates the platform, tries different inputs and reports bugs for human QA reviewers. Hidalgo views this as the foundation for a new software development lifecycle, with many cloud agent sessions testing features in parallel.
*The complete study is available on the* [*Asana*(opens in a new window)](https://asana.com/inside-asana/cut-browsers-agent-cost) *and* [*StackAI*(opens in a new window)](https://www.stackai.com/blog/how-stackai-by-asana-used-astra-codex-and-command-to-reduce-browser-agent-costs) *blogs.*
## Join the new era of work
More than 1 million businesses around the world are achieving meaningful results with OpenAI.
[Contact sales](https://openai.com/contact-sales/)
## Keep reading
Oracle — Cover
[How Oracle turns days of work into minutes with ChatGPT and Codex
Oct 8, 2026](https://openai.com/index/oracle/)
1Password > Card image > Fiber ridge close-up
[1Password increases engineering productivity 21% with Codex
Sep 8, 2026](https://openai.com/index/1password/)
loveholidays customer story art card v4
[How loveholidays is making everyone a builder with Codex
Aug 26, 2026](https://openai.com/index/loveholidays/) | OpenAI | — | — |
2026-10-10T22:31:58Z
Case continues cooling in deep tail (~276h, 0 pts/h) with static three-platform periphery. The DevDay article re-fetch is confirmatory context (same launch-day content), not a material shift. Asana case study (vendor-published, 76x browser-agent cost reduction) was already incorporated and conflates model switch with caching/history-engineering gains. All seven independent API-side measurement lines stand; subscription-side quota burn contradiction stands; capability breaks on creative work stand; agent-loop turn-inflation anecdotes (2) remain unreplicated; independent coding-agent-harness cost/capability results, adoption-at-scale data, and systematic turn-count evidence remain absent. The '(yet)' hedge on the sole Codex-harness speed measurement persists. Microsoft weights rumor unverified after 3+ days.
2026-10-10T06:53:05Z
Asana case study (vendor-published) adds first enterprise browser-agent data point showing 76x cost reduction and 5x speedup with GPT-6.1 Sol in Codex, but conflates model switch with caching/history-engineering gains. Core unresolved branches persist: systematic turn-inflation on non-browser agent loops, hidden multipliers, adoption at scale, independent coding-agent-harness cost/capability results, and Codex speedup replication. Case continues cooling in deep tail (~260h, 0 pts/h) with static three-platform periphery.
2026-10-10T03:23:15Z
Asana case study (vendor-published) reports 76x cost reduction and 5x speedup on browser-agent workloads using GPT-6.1 Sol in Codex with workflow optimization (caching, history budget, batch screenshot removal). This is the first enterprise production data point directly on agent-workload economics, but it conflates model switch with engineering optimization and remains a single OpenAI-authored study. The core unresolved branches persist: systematic turn-inflation on non-browser agent loops, hidden multipliers, adoption at scale, independent coding-agent-harness cost/capability results, and Codex speedup replication.
2026-10-09T19:58:40Z
evidence attached: openai.article.10dcfe0bfbc53363824c075e — Asana's reported 76x cost reduction using GPT-6.1 Sol in Codex browser-agent tests directly bears on whether Sol becomes the cost-efficient default for agent workloads.
2026-10-09T07:14:47Z
Case continues cooling in deep tail (~237h, 1.67 pts/h) with static three-platform periphery. New evidence (1x17pdr: GPT-6 interactive UI rollout) is confirmatory context aligned with the Sol launch, not a material shift on any resolution branch. The magnitude-valve reading reflects launch-day concentration in existing communities, not expanding periphery. All seven independent API-side measurement lines stand; subscription-side quota burn contradiction stands; capability breaks on creative work stand; agent-loop turn-inflation anecdotes remain unreplicated; independent coding-agent-harness cost/capability results, adoption-at-scale data, and systematic turn-count evidence remain absent. The '(yet)' hedge on the sole Codex-harness speed measurement persists.
2026-10-09T04:52:45Z
evidence attached: reddit.post.1x17pdr — Rollout of GPT-6 with interactive UI (charts, forms, tools) aligns with the GPT-6.1 Sol launch and its agent-oriented feature set.
2026-10-08T08:58:22Z
grounded: converges/high — OpenAI has productized the exact scout–senior / prefix-caching economics Scott's canon argues for: near-flagship capability at 1/5 token price with $0.10/Mtok c
2026-10-08T08:45:23Z
New anecdotal report (1x0l5bp) of 6x-Sol quality failures driving a Claude Code Opus 5.5 switch reinforces the existing capability-break pattern on unmeasured/creative work, but does not alter the settled API-yes / subscription-no / capability-below-flagship split. The Microsoft weights rumor remains unsourced after ~3 days with no cross-platform pickup. The independent Codex-harness speed measurement (1wzzwt2) is n=1 with a '(yet)' hedge on rollout lag. Episode at ~214h in deep tail: 0 pts/h, static 3-platform periphery, same unresolved branches (agent-loop turn inflation systematic?, hidden multipliers?, adoption at scale?, independent coding-agent-harness cost/capability results?).
2026-10-08T08:40:10Z
evidence attached: reddit.post.1x0l5bp — User reports 6x-Sol quality failures and Astra cost-prohibitive, switching to Claude Code Opus 5.5 — direct hands-on evidence bearing on whether Sol becomes cost-efficient default for agent workloads or underperforms Astra.
2026-10-07T17:15:25Z
First independent coding-agent-harness measurement arrives — and it contradicts the vendor: a controlled frozen-benchmark rerun (hash-checked prompt, hidden tests, median of 3–5 runs, Opus 5.5 control) finds none of OpenAI's claimed ~50% Codex speedup for Sol, with Astra's median slowed — adding a claims-vs-measurement gap and a live check on the 28-day rollout, without touching the settled API-yes/subscription-no split. Low heat holds despite magnitude-valve and 'accelerating' momentum readings: 1.33 pts/h at ~199h is ~0.2% of peak, the momentum figure is a near-zero-base artifact of the new post, the periphery is still three static platforms, and the 695-pt Microsoft weights rumor remains unsourced with no cross-platform pickup after ~3 days.
2026-10-07T16:35:56Z
evidence attached: reddit.post.1wzzwt2 — Independent frozen-benchmark rerun fails to reproduce OpenAI's claimed ~50% Codex speedup for Sol (and Astra slowed), directly bearing on whether Sol's cost-efficiency claims hold.
2026-10-06T12:22:52Z
The tail 're-acceleration' (13.5→59 pts/h, 94th percentile, accelerating) decomposes to the same unsourced Microsoft weights-rumor post climbing alone (163 pts/h vs p90 132) — its comments are still link-requesting and skeptical, no derivatives, no cross-platform pickup — and the Codex-vanishing report failed to replicate (0 pts, 0.5 ratio, dying), so the case's meaning is unchanged. Low heat holds over the magnitude-valve and percentile readings because the periphery remains 3 static platforms with the top-decile objects all from the launch window; a sourced leak pickup or replicated rollout blip stays the re-heat trigger.
2026-10-06T11:33:25Z
evidence attached: reddit.post.1wyy8fx — Single-user report that 6.1-sol vanished from Codex directly bears on whether Sol holds rollout and becomes the cost-efficient default — weak but case-resolving evidence to watch for replication.
2026-10-06T05:56:55Z
New tail evidence is an unsourced Reddit claim that Microsoft leaked 6.1 Sol as the failed GPT-6 Sol base weights run at 2 inference passes instead of 3 — if corroborated it recasts 6.1 as a serving-config retune (revising the rebranded-Astra-Minor backstory and sharpening the model-perishability angle), but it touches no resolution branch since every measured line was run on 6.1 as shipped, so the settled API-yes/subscription-no split stands. Not material pending the Microsoft doc or independent corroboration; the 13.5 pts/h uptick (91st percentile, 'accelerating') is one rumor post moving inside a static 3-platform periphery — low heat holds for the same launch-day-concentration reason as before, and a sourced or cross-platform pickup of the leak is the stated re-heat trigger.
2026-10-06T05:26:22Z
evidence attached: reddit.post.1wyu0ss — Claim that Sol is the same base weights run with fewer inference passes directly bears on what Sol is and its cost-efficiency story, which is exactly what resolves that case.
2026-10-06T01:52:17Z
The new scraping-cost report is a second, workload-distinct datapoint for agent-loop cost inversion (Sol 6.1 burned 40% of a weekly quota on an unfinished brand where Opus 5.5 cleared ~80) and anecdotal confirmation of the already-settled subscription-burn contradiction — it thickens an existing branch rather than changing what the case means, so not material; the 3.5 pts/h 'spike' is a 10.5x multiple on a near-zero baseline, and low heat holds over the magnitude-valve flag because the top-decile spread remains launch-day concentration inside a static 3-platform periphery with no new implementations or outlets.
2026-10-05T23:34:20Z
evidence attached: reddit.post.1wylyxw — Early hands-on cost data against Sol's efficiency pitch: one unfinished brand burned 40% of a weekly limit that let Opus 5.5 scrape 80 brands — exactly the underperformance evidence the case's resolution tracks.
2026-10-05T22:06:40Z
The re-fired sensor is a mis-scoped attach: the new post is a 5.6-Sol laziness complaint (0 score, 45% upvote ratio), not 6.1 evidence, and its comments merely restate the known effort-depth tradeoff — tail noise, not a meaning change. The settled split (API-yes / subscription-no vs Opus 5.5 / below-Astra on unmeasured work / unreplicated over-checking anecdote) stands, with adoption-at-scale and agent-loop turn-count inflation as the only open axes and the episode in a dead 0.5 pts/h tail.
2026-10-05T20:39:22Z
evidence attached: reddit.post.1wyhv7s — Early hands-on report of Sol underdelivering answers is a small datapoint for the case's criterion that early quality reports decide whether Sol holds as the cost-efficient default.
2026-10-05T07:28:27Z
No meaning change since the over-checking reprice: the re-fired velocity_spikes are score jitter on the pinned FrontierMath thread (575↔580 across three events, six days old) plus one-comment accrual on two minor threads — tail noise, not velocity. The settled split stands (API-yes / subscription-no vs Opus 5.5 / below-Astra on unmeasured work / over-checking single anecdote with visible community pushback), with adoption-at-scale and agent-loop turn-count inflation as the only open axes; low heat holds because the magnitude-valve spread is launch-day concentration inside three static communities, not expanding periphery, and the episode is at 0.0 pts/h against a 398 peak.
2026-10-04T13:30:17Z
The over-checking report adds a qualitatively new failure mode: on an actual agent loop, 6.1 burned ~18x Astra's wall-time in redundant verification, so the one-fifth token price can invert exactly on agentic workloads where cost = turns × tokens — the workload class the hypothesis and Scott's routing decisions hinge on — sharpening the open cost-efficiency question from subscription burn to turn-count inflation. The release audit stays settled (API-yes / subscription-no / below-Astra on unmeasured work), adoption-at-scale remains the sole open axis, and attention is now dead (0.0 pts/h tail; the velocity_spike was late accrual on the 5-day-old FrontierMath thread).
2026-10-04T13:25:59Z
evidence attached: reddit.post.1wxfc4r — Early hands-on report of the 6.1-generation model pathologically over-checking on an agentic task (~1.5 days vs 1-2 hours for Astra/Sol) — direct counter-evidence to whether the 6.1 line becomes the cost-efficient agent-workload default.
2026-10-03T20:03:16Z
The $100-plan quota anecdote (2-hour job = 2–3% of weekly quota, user eyeing a downgrade) consolidates the cost-efficiency side of the already-settled split without flipping anything: API-yes / subscription-no vs Opus 5.5 / below-flagship on unmeasured work stands, and adoption-at-scale remains the sole open resolution axis. Low heat holds despite the magnitude-valve flag and 81st-percentile peer reading — the absolute rate is ~2.7 pts/h against a 398 peak, momentum cooling, and the top-decile spread is four-day-old launch residue inside three static communities, not expanding periphery.
2026-10-03T19:26:18Z
evidence attached: reddit.post.1wwur9i — On-hypothesis real-usage economics: a 2-hour coding job consuming 2–3% of weekly quota supports Sol's cost-efficiency side of the resolution.
2026-10-03T17:44:03Z
grounded: converges/high — OpenAI is now selling as productized tiers — near-Astra capability at $2/$10 per Mtok with $0.10/Mtok cached input, corroborated by seven independent measured l
2026-10-03T17:36:18Z
The Open Pro defection report (Sol 6/6.1 refuses detailed plans, quota exhausted post-reset, user trades for Opus 5.5 resets) consolidates the subscription-side contradiction with a concrete symptom but flips nothing: the release-claim audit is now settled — API-cheap yes, subscription-no, capability below Astra on unmeasured/structured work — leaving adoption-at-scale as the sole open resolution axis. Attention is pure tail accrual (5 pts/h vs 389 peak, cooling, static 3-platform periphery); low heat holds and the magnitude-valve reading remains launch-day concentration within static communities.
2026-10-03T17:23:53Z
evidence attached: reddit.post.1wwshhj — Early hands-on report that Sol 6/6.1 refuse detailed plans and the user defected to Claude — exactly the underperformance evidence the case resolves on.
2026-10-02T18:47:45Z
Epoch-linked FrontierMath Tier 4 100% makes third-party math a second bracketing line (with MathArena #1) showing Sol Max at/above Astra on measured math — consolidating, not flipping, the settled split: harness-verifiable work ≈ flagship at ~1/5 price, unmeasured creative/building work below flagship, subscription still not cheaper than Opus 5.5. Not material under the strict bar: same domain and direction as the existing MathArena line, and the episode remains tail accrual (2.2 pts/h vs 272 peak, static 3-platform periphery, no new implementations), so low heat holds despite the magnitude-valve flag; adoption-at-scale is the only open resolution axis.
2026-10-02T17:39:07Z
evidence attached: reddit.post.1wvzs2k — Epoch-linked 100% FrontierMath Tier 4 result is independent benchmark capability evidence directly bearing on whether Sol's release claims hold against the underperformance counter-signals.
2026-10-02T11:25:14Z
A sixth independent measured line — a deterministic photo-to-Blender agent-loop benchmark scoring Sol 61 vs Astra 66 at ~1/10 the cost ($0.36 vs $3.91) — quantifies the Blender capability discount as modest (~5 pts) on the exact domain where anecdote reported failure, consolidating the pattern: small measured discount, large price win on harness-verifiable work. No branch flips; adoption-at-scale remains the open resolution axis, and the residual engagement is tail accrual on two Reddit threads rather than periphery expansion, so low heat holds despite the magnitude-valve flag.
2026-10-02T11:23:41Z
evidence attached: reddit.post.1wvqfqe — Early hands-on agentic benchmark with deterministic scoring shows Sol at 61 vs Astra 66 for ~1/10th the cost — exactly the cost-efficiency-vs-capability evidence the Sol case resolves on.
2026-10-01T13:27:08Z
Two more low-engagement same-platform hands-ons add adoption texture without flipping a branch: a Plus-tier report of ~5% vs Astra's 20–30% credit drain supports cheaper-than-Astra at the subscription tier (consistent with, not overturning, the controlled Opus-5.5 quota contradiction), and a 34-pt consensus thread nets 'real improvement over garbage 6-Sol, below Astra, slow on long tasks' with users still defaulting to Astra. The settled API-yes / subscription-vs-Opus-no / below-flagship-on-unmeasured-work split stands with adoption unobserved at scale; the episode keeps cooling (~9 pts/h vs 241 peak, static 3-platform periphery), so the 93rd-percentile peer reading is launch-day concentration, not expansion, and heat drops to low pending an event-driven re-trigger.
2026-10-01T13:22:48Z
evidence attached: reddit.post.1wuyrq9 — Hands-on Plus-tier data point: Sol 6.1 drains ~5% vs Astra's 20-30% of usage credit on routine work, directly supporting the cost-efficient-default side of the open case.
2026-10-01T12:24:02Z
evidence attached: reddit.post.1wuxyo2 — Early independent hands-on says 6.1 improved over 6 Sol though the user still defaulted to Astra — direct evidence on the case's stated resolution axis.
2026-10-01T10:45:10Z
The 'material first-party article revision' is a refetch of the DevDay 2026 Recap whose retrieved text adds no new Sol claims, pricing, or rollout detail beyond the existing Keep reading link — housekeeping, not a meaning change. The case holds its settled API-yes / subscription-no / capability-below-flagship split with adoption unobserved, and the episode continues cooling (~2.3 pts/h vs 196 peak, static 3-platform periphery), so low heat stands despite the magnitude-valve flag: the top-decile spread is launch-day concentration within static communities, not expansion.
2026-10-01T09:37:27Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-10-01T09:23:35Z
evidence attached: reddit.post.1wuurt6 — Builder account of GPT 6 Sol replaced by 6.1 Sol after seven days, with pricing and eval-churn detail bearing on whether Sol becomes the cost-efficient default.
2026-09-30T23:36:26Z
CafeBench (value-strong, below Opus) is a fifth independent line landing squarely inside the already-settled API-yes/subscription-no/capability-below-flagship split — confirmatory, not meaning-changing. The launch episode has passed its information peak: periphery fixed at the same 3 platforms since launch, no new implementations/outlets/communities, momentum cooling, newest item 5 points. I price heat low despite the magnitude-valve reading because the top-decile spread is launch-day concentration and thread depth within static communities, not expanding periphery, and the open questions (production coding-agent evidence, adoption at scale, hidden multipliers) are weeks-scale and best carried by event triggers rather than clock cadence.
2026-09-30T21:37:57Z
evidence attached: hn.story.49913486 — Third-party CafeBench result (strong value, below Opus tier) is direct evidence on whether Sol becomes the cost-efficient agent default or underperforms the frontier.
2026-09-30T20:25:09Z
The four attached lines split the cost branch rather than settling it: independent measurements (MathArena #1 at 86.3%/$0.94 vs Astra 81.9%/$2.26, the AA index, the PokeBench record, and 3.4x cheaper per task at API prices) now corroborate near-Astra-at-1/5-price on measured/API workloads, while the same controlled re-test contradicts the subscription pitch (~38% more quota burn than Opus 5.5 at equal intelligence). The case's meaning shifts from 'does Sol match Astra' to an API-yes/subscription-no cost-default split still awaiting adoption and coding-agent evidence; the velocity spike is one late cost-positive thread, not spread, so medium heat holds despite the magnitude-valve reading.
2026-09-30T18:42:40Z
evidence attached: reddit.post.1wub1zx — Controlled re-test measuring Sol 6.1 burning ~38% more subscription quota than Opus 5.5 despite API-price advantage is direct third-party evidence on whether Sol becomes the cost-efficient agent default.
2026-09-30T18:42:39Z
evidence attached: reddit.post.1wu8zb2 — Independent MathArena results (86.3% at $0.94 vs Astra 81.9% at $2.26) are third-party corroboration directly bearing on whether Sol becomes the cost-efficient default.
2026-09-30T18:42:39Z
evidence attached: reddit.post.1wuatc8 — First-hand Claude Max user seriously weighing defection to Sol on per-task cost, direct adoption-side evidence for the Sol-as-default question.
2026-09-30T18:42:39Z
evidence attached: reddit.post.1wu854w — Documents the 7-day lifespan of GPT-6-SOL and community verdict it underperformed 5.6, material rollout context for whether Sol becomes the cost-efficient default.
2026-09-30T13:37:34Z
A fourth independent user line adds the first direct user-side support for the cost-efficiency branch — similar work-problem throughput to Astra at dramatically lower token cost — counterweighting the usage-burn signal and splitting general dev/work by cost-vs-capability priority rather than a clean hold-or-break. With momentum flipped to cooling (23 pts/h vs a 106 peak) and periphery fixed at three platforms, this remains corroborated depth, not spread, so heat holds at medium.
2026-09-30T13:26:03Z
evidence attached: reddit.post.1wu4fxh — Early hands-on report directly supporting the cost-efficiency branch of the case: 6.1 Sol completing similar work at dramatically lower token cost than Astra.
2026-09-30T11:50:43Z
A third independent user line places 6.1 between Sol and Astra on ordinary dev work, with long-task slowness and usage burn giving the first weak user-side signal on the quota question — shifting the case's shape from 'holds on measured tasks, breaks on creative work' to a measurable index-vs-experience gap at 1/5 price. The decisive unknowns (coding-agent-class evidence, adoption) are unchanged, and the newest object is itself thin (1 pt), confirming depth-not-periphery, so heat holds at medium.
2026-09-30T11:24:57Z
evidence attached: reddit.post.1wu1zx7 — Early hands-on placing 6.1 between Sol and Astra with long-task slowness and usage burn — first user-side evidence on whether 6.1 lives up to its positioning.
2026-09-30T10:45:06Z
The only new object is a duplicate HN submission of the already-counted Artificial Analysis finding (3 pts, died on arrival) — not a new measurement line; comment growth on the two Reddit threads reiterates the recorded workaround consensus rather than adding a workload class or implementation. The two-sided interpretation stands unchanged: the speedometer's acceleration (19 pts/h, 91.5th percentile) is thread depth on one platform, not widening periphery, so heat holds at medium with event-driven re-triggers on coding-agent benchmarks, production hands-on, or quota/pricing reports.
2026-09-30T10:24:59Z
evidence attached: hn.story.49906669 — shared external link with case evidence
2026-09-30T05:58:24Z
The vendor's core claim has crossed the evidence-class ladder: Artificial Analysis's independent index reportedly slots Sol 6.1 near Astra (replacing GPT-6 Sol within a week), joining the judge-free PokeBench run, so the 1/5-price capability claim now rests on two independent measurement lines rather than OpenAI-run evals alone. The live contest narrows to workload-class generalization — the Blender regression thread is now the case's most-discussed object with the community converging on cheap-model multi-pass workarounds — while coding-agent evidence, the class that would re-point Scott's routing defaults, is still absent.
2026-09-30T05:26:12Z
evidence attached: hn.story.49900328 — Independent Artificial Analysis measurement of GPT-6.1 Sol replacing GPT-6 Sol in 7 days at near-Astra intelligence is exactly the third-party performance evidence the case is waiting on.
2026-09-30T04:32:48Z
The case has split into a genuine two-sided evidence contest: the first independent same-harness measurement (PokeBench — Sol 6.1 beats Astra's record, 244 vs 246 turns, at $2.60 vs $13.49) confirms the one-fifth-price claim on computer-use, while two hands-on reports (Blender game-dev, browser polish) show Sol trailing Astra badly on creative/building work. The open question shifts from 'is the vendor claim real?' to 'on which workload classes does the trade hold?' — coding-agent evidence, the part that bears on Scott's routing defaults, is still absent.
2026-09-30T03:31:24Z
evidence attached: reddit.post.1wtrf1x — Independent measured benchmark result showing GPT-6.1 Sol beating Astra's PokeBench record at ~1/5 cost with same-harness methodology — direct evidence on the case's cost-efficiency-vs-Astra question.
2026-09-30T03:31:24Z
evidence attached: reddit.post.1wttjzi — Early hands-on report of Sol 6.1 Max badly underperforming on real game-dev work — exactly the underperforming-Astra evidence the case waits on.
2026-09-30T01:08:06Z
grounded: converges/high — OpenAI is now selling the exact input economics Scott's scout-and-senior split and Prefix-Caching Economics argued for — near-flagship capability at one-fifth p
2026-09-30T00:59:49Z
case created — A distinct new frontier release with a resolvable price-performance claim, anchored on OpenAI's own DevDay recap, already carrying a contrarian hands-on report.