2026-10-11 17:14 UTC

OpenAI claims its launched GPT-6.1 Sol — rolling out across ChatGPT Work, Codex, and the API alongside a new $500/month Pro tier — approaches Astra-level coding and computer-use performance at one-fifth the token price; whether Sol actually becomes the cost-efficient default for agent workloads, or early hands-on reports of it underperforming Astra hold, resolves it.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highgpt-6-1-sol inference-economics openai coding-agentsOpenAI
Surfaced 2026-10-01T09:40:27Z — DevDay 2026 Recap — The new seventh-day-churn post re-quotes vendor-run evals (DeepSWE parity with Astra, AutomationBench-vs-Opus at ~1/3 price) already captured at launch and adds the same 7-day rebrand narrative — confirmatory context, not a meaning change: the case remains the settled API-yes / subscription-no / capability-below-flagship split with adoption unobserved. The episode keeps cooling inside the same three platforms; low heat holds despite the magnitude-valve reading because the top-decile spread is launch-day concentration within static communities, not expanding periphery.

What is this?

OpenAI launched GPT-6.1 Sol at DevDay on September 29, 2026 — a week after the short-lived GPT-6 Sol — positioning it as a near-Astra model for agentic coding, computer use, and professional work at roughly one-fifth Astra's token price ($2/$10 per Mtok input/output, $0.10/Mtok cached input). The model rolled out same-day in ChatGPT Work, Codex, and the API alongside a new $500/month ChatGPT Pro tier. OpenAI's own benchmarks claim Sol matches Astra on DeepSWE v1.1 at ~1/5 cost, reaches 71.4% on OSWorld 2.0 (vs Astra's 73.5%) at ~1/7 cost per task, and improves on AutomationBench. Independent third-party measurements (PokeBench, Artificial Analysis, MathArena, CafeBench, deterministic photo-to-Blender loops, FrontierMath Tier 4) largely corroborate near-Astra capability at dramatically lower API-side cost on harness-verifiable tasks. However, the same controlled tests contradict the subscription value pitch: Sol 6.1 burns ~38% more $20-plan quota than Opus 5.5 at equal intelligence, and hands-on reports show capability breaks on unmeasured/creative work (Blender game assets, detailed planning refusals) plus two anecdotes of agent-loop turn-count inflation that could invert the token-price advantage exactly on the agent workloads the hypothesis targets. Adoption at scale and independent coding-agent-harness cost/capability results remain unobserved.

Why it matters to Scott

OpenAI has productized the exact scout–senior / prefix-caching economics Scott's canon argues for: near-flagship capability at 1/5 token price with $0.10/Mtok cached input, corroborated by seven independent measurement lines on harness-verifiable work. The case provides the dated-receipts moment — a consequential vendor independently arriving at the architecture Scott built (model barbell, scout-senior split, prefix-caching economics) — while simultaneously surfacing new failure modes that bear on his routing layer: subscription-side quota burn contradicting the API economics, agent-loop turn-count inflation that could invert cost = turns × tokens on the exact workload class his LiteLLM cheap/medium aliases and plan-forecast routing target, and a vendor-claim vs independent measurement gap on Codex speed. Adoption at scale and independent coding-agent-harness cost/capability results remain the open resolution axes.
ip:framework.scout-senior-splitip:concept.model-barbellip:concept.prefix-caching-economicsip:concept.token-economicsip:concept.ai-unit-economicsip:framework.agent-loopip:concept.agent-receiptsip:concept.agent-observabilityip:concept.model-perishabilityip:framework.nuke-and-regeneratedev:technology.litellmdev:concept.plan-forecast-model-routingdev:concept.cost-tiered-llm-routingdev:project.askdev:concept.propose-finalise-gateradar:anthropic-opus55-cache-read-repricingradar:hidden-reasoning-real-task-costsradar:agent-run-cost-unpredictabilityradar:chatgpt-pro-200-signup-pauseradar:codepress-subscription-cloud-agentsradar:claude-sonnet-5-permanent-pricingradar:bytedance-seed-2-code-validationradar:agent-bottling-benchmarkradar:llama-cpp-byte-identical-prompt-dedupradar:fractal-recursive-agent-loops
queries asked of Scott's wikis
  • inference-economics prefix-caching economics scout-senior split
  • open-weights strategy model sovereignty local inference
  • coding-agent routing defaults LiteLLM cheap medium aliases
  • agent-loop turn-count inflation hidden multipliers
  • frontier-model perishability 7-day churn pattern
  • subscription vs API pricing divergence quota burn

Measured heat

now 0 pts/hpeak 566 pts/hcomments 0/hpeers p15momentum: steady3 platformsage 294h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-29 10:00⭐ origin directly observedDevDay 2026 Recap
OpenAI on openai
—
09-29 20:37first on r/OpenAI · published · +10.6hFirst impression about gpt6.1 sol
Individual_Art_5163
—
09-29 20:52first on hacker news · published · +10.9hArtifical Analysis - GPT-6.1 Sol Replaces 6 Sol After 7 Days
Fe2O3
—
09-30 16:08first on r/singularity · published · +30.1hGPT-6.1 Sol is now #1 on MathArena: 86.3% accuracy for $0.94, beating Astra’s 81.9% at $2.26
141_1337
—
09-30 17:16first on r/ClaudeAI · published · +31.3hWatched the DevDay stuff yesterday and for the first time I'm actually rethinking my $200 Claude Max sub.
OccasionNo4703
—
10-09 00:46first on r/artificial · published · +230.8hOpenAI rolls out GPT-6 with Intelligent UI — ChatGPT answers can now include interactive tools
winer666
—
10-09 07:00first on openai · published · +237.0hAsana cuts model costs 76x in browser tests with GPT-6.1 Sol
OpenAI
—
09-29 20:37amplified on r/OpenAIreddit.post.1wtlf5y
Individual_Art_5163
peak 3 · 1 comments · 0% of case engagement
09-29 20:52amplified on hacker newshn.story.49900328
Fe2O3
peak 2 · 0 comments · 0% of case engagement
09-29 20:56amplified on r/OpenAIreddit.post.1wtlwy4
ebytes111
peak 3 · 3 comments · 0% of case engagement
09-30 00:56amplified on r/OpenAIreddit.post.1wtrf1x
VibeCodyH
peak 157 · 18 comments · 5% of case engagement
09-30 02:40amplified on r/OpenAIreddit.post.1wttjzi
Comfortable-Cat-9611
peak 80 · 81 comments · 5% of case engagement
09-30 09:59amplified on hacker newshn.story.49906669
theanonymousone
peak 80 · 99 comments · 10% of case engagement
22 more amplifiers in ainews.case_chain
09-30 00:20our radar first saw it · +14.3hdiscovery anchor: openai.article.c6a404b23739c6d6a0744de1—
10-01 09:37reached heat=high · +47.6h · via ledger——
pace: p96 vs 1188 stories at the 168h mark (now 294h old) — ahead of sonnet-55-release-economics (1.0x), behind griffin-video-turing-test-claim (1.0x)

Evidence (30) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 openai ⭐DevDay 2026 Recap
Retrieved article excerpt

Open article · Retrieved 2026-10-10T21:31:18.902689+00:00

September 29, 2026

[Company](https://openai.com/news/company-announcements/)[Product](https://openai.com/news/product-releases/)

# DevDay 2026 Recap

At DevDay 2026, we’re giving people more ways to take on ambitious work and tools to build what comes next.

Share

DevDay 2026 is our biggest yet, with more than 20 major announcements across ChatGPT, Codex, our models, and entirely new forms of working with AI.

We believe AI can help bring about a new renaissance of creativity and discovery. It should give people more time for what matters to them, more freedom to pursue their ideas, and the ability to do things they didn’t think were possible.

Today, we introduced agents that can take on ongoing responsibilities and new ways for people and AI to work together. We also expanded our commitment to an open ecosystem by opening up ChatGPT as a shared surface where humans and agents can collaborate and where developers can directly launch new native experiences to our collective 1.2B weekly users. 

Here’s everything we announced.

[New ways of working](https://openai.com/index/devday-2026-recap/#a-whole-new-way-to-work-with-ai)[Build with Codex & the API](https://openai.com/index/devday-2026-recap/#new-tools-for-developers-in-codex-api)[Customize ChatGPT with Plugins](https://openai.com/index/devday-2026-recap/#more-customization-in-chatgpt-with-plugins)[People & AI working together](https://openai.com/index/devday-2026-recap/#people-and-ai-working-together)[Do more with your ChatGPT subscription](https://openai.com/index/devday-2026-recap/#do-more-with-your-chatgpt-subscription)

## A whole new way to work with AI

GPT-6.1 Sol title on a dark, starry background.

Astra Ultrafast title over a starfield, with availability in ChatGPT, Codex, and the API.

Screenshot of a project-specific policy table showing Zero Data Retention with Private Safety Processing and validated external storage.

### Dots

Dots are remarkably capable, always-on agents built to handle everything.They’re a whole new way to work with AI—one that gets to know what matters to you, is always working on your behalf, and takes important work off your plate so you get more of your time and attention back.

Available on Pro and Business Premium in eligible markets. Enterprise, Edu, and Healthcare users can try the beta when their workspace admin enables it; it is off by default. [Learn more⁠](https://openai.com/index/introducing-dots/)

### GPT-6.1 Sol

We’re introducing GPT‑6.1 Sol, a major upgrade to GPT‑6 Sol with exceptionally strong performance on agentic coding, along with computer use, and professional work. It delivers near-Astra intelligence to everyone at a fifth of its standard input and output token prices, giving developers more flexibility to execute complex workflows within the same budget.

Available to all API, Plus, Pro, Business, Enterprise, and Edu users. [Learn more⁠](https://openai.com/index/introducing-gpt-6-1-sol/)

### Ultrafast

Ultrafast is our premium speed tier for workloads where speed matters most. Ultrafast offers up to 8× faster token generation (300 tokens per second) in Codex and up to 6x in the API.

GPT‑6 Astra Ultrafast is available today in the [API⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/ultrafast-mode) and in [ChatGPT Work and Codex⁠(opens in a new window)](https://learn.chatgpt.com/docs/agent-configuration/speed) on Pro 500 and Enterprise plans. GPT‑6.1 Sol Ultrafast is coming soon.

### Private Intelligence

OpenAI Private Intelligence helps businesses use frontier AI with greater confidence that their data is protected. [Zero Data Retention with Private Safety Processing⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/private-safety-processing) enables automated safety reviews without giving OpenAI personnel access to the underlying content. Our preview of Private Inference, coming this fall, combines confidential computing with strict, verifiable controls.

[Contact us⁠](https://openai.com/form/private-intelligence-interest/) about Private Intelligence

## New tools for developers in Codex & the API

Codex cloud projects interface with a purple cloud icon.

Codex CLI terminal interface.

Codex code review interface with a discussion of a code change.

Codex Security Cloud findings dashboard.

Decisions API text on space background.

Diagram of Agents API.

OpenAI and AWS logos on a blue and purple gradient.

### Codex in the cloud

Now developers can run Codex wherever they need it: on a computer, remotely from a phone, or in the cloud from any device. Reusable development environments help tasks start quickly and give your team a shared setup with approved settings and permissions.

Available on Plus, Pro, Business, Healthcare, Education, and Enterprise. [Learn more⁠(opens in a new window)](https://learn.chatgpt.com/docs/cloud)

### A refreshed Codex CLI

The Codex CLI now lets you start and steer tasks with your voice. The new /agents view makes it easier to delegate work and track multiple tasks at once. We’ve also improved everyday workflows to help  you edit prompts, resume sessions, and use worktrees, while a cleaner terminal UI makes longer sessions easier to read.

Available to all plans. [Learn more⁠(opens in a new window)](https://learn.chatgpt.com/docs/codex/cli)

### Code Review

The new code review experience in the ChatGPT desktop app makes it easier to review changes across projects. You can read summaries, explore diffs, and ask Codex about potential issues before sharing feedback on GitHub pull requests or GitLab merge requests. With automatic reviews, Codex can also take a first pass in the cloud while you’re away.

Available on all plans. [Learn more⁠(opens in a new window)](https://learn.chatgpt.com/docs/code-review?surface=app)

### Codex Security Cloud

Codex Security Cloud gives defenders a better set of tools to harden their infrastructure. Scan entire GitHub repositories on demand or on a schedule, with ongoing checks of new commits. Codex investigates findings, removes duplicates and prepares fixes in the cloud, even with your laptop closed. It includes access to models offered through [Daybreak Blue⁠](https://openai.com/daybreak/) without a separate Daybreak application.

Available to all Pro, Business, Enterprise and Edu users on desktop and web. [Learn more⁠(opens in a new window)](https://learn.chatgpt.com/docs/security/setup)

### Decisions API

Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent’s next action.

Available in limited preview today with a broad release planned in the coming days.

### Agents API with Computer use

The [Agents API⁠](https://openai.com/index/introducing-the-agents-api/) now supports computer use so developers can build agents that interact with software to complete tasks. It also brings Codex’s multi-agent capabilities, tool search, tool calling, and context compaction into your application. OpenAI runs the underlying infrastructure so your team can focus on building the application.

Available through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise. [Learn more⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/agents-api/tools/computer-use)

### Bedrock Managed Agents, powered by OpenAI

OpenAI worked with Amazon on Bedrock Managed Agents to take the core capabilities of the Agents API and add customization to work natively in AWS and integrate with AWS resources. Now you can use Bedrock Managed Agents to build OpenAI agents that run entirely in AWS. [Learn more⁠(opens in a new window)](https://aws.amazon.com/bedrock/managed-agents-openai/)

## More customization in ChatGPT with plugins

Canva plugin extension shown in the ChatGPT sidebar.

Figma, Shopify, and plugin cards on a blue background.

Three example websites created with Sites.

Diagram linking MCP Events with automations.

### Plugin extensions

We’re opening the platform we use to build ChatGPT features so developers can create their own experiences within ChatGPT. Plugin extensions let you give your plugin a home in the sidebar and build interactive panels where people can work alongside the conversation. You can also create viewers for the file types your product supports.

Available to all plans. [Learn more⁠(opens in a new window)](https://developers.openai.com/plugins/build/extensions)

### Improved plugin creation, submission, and discovery

We’re making it easier to build plugins and help people discover them. Plugin Creator helps you build your plugin while a redesigned submission flow provides clearer feedback. Improved ranking and recommendations help people find relevant plugins in the directory and in conversations. Users choose which plugins to use and approve the access each one receives.

Available to all plans. [Learn more⁠(opens in a new window)](https://developers.openai.com/plugins)

### Sites can now host plugins

You can now add supported ChatGPT plugins to the Sites you build. Teammates in your workspace can use the same app with their own connected data and permissions. We’re also making automations easier to add and manage so Sites can keep shared information up to date.

Available to Business, Enterprise, Healthcare, and Edu plans. [Learn more⁠(opens in a new window)](http://chatgpt.com/features/sites/)

### MCP events for plugin automations

We’re adding support for the [proposed MCP Events specification⁠(opens in a new window)](https://modelcontextprotocol.io/community/working-groups/triggers-events), so plugins can start automations when something happens in a connected app. For example, you can ask ChatGPT to watch for new tasks on a project board. When one comes in, ChatGPT can read the linked documents and draft a plan, even while you’re away.

Available to all plans. [Learn more⁠(opens in a new window)](https://developers.openai.com/plugins/build/mcp-events)

## Improving how people and AI work together

Array of ChatGPT Space windows in a starry field.

Collaborative ChatGPT Page for Q4 event planning with edits from teammates.

Collaborative presentation editor with slide previews and comments.

Schedule a team task interface with team selection.

Microsoft Teams and Slack icons above an @ChatGPT message.

Meeting recap with an @ChatGPT action request.

Shareable profile for Bailey with activity and showcased creations.

### ChatGPT Space

A new home for your team to collaborate with AI to get work done. Create a dedicated space where teammates, ChatGPT, and your dot can build on shared knowledge. ChatGPT can keep your space organized based on instructions you provide, so you can quickly find what you need and pick up where you left off.

Available to all Pro, Business, and Enterprise plans on the ChatGPT desktop app and web, with finding, reading, and sharing pages also available on mobile; mobile creation and editing are coming soon. [Learn more⁠(opens in a new window)](https://chatgpt.com/features/space/)

### Pages

Pages are a new type of document, built for human and agent collaboration. Anything you can do in ChatGPT you can do on a page: write, research, generate charts, create images, or visualize information. Create a page in a conversation, then invite your team to contribute ideas and feedback.

Available to all Pro, Business, and Enterprise plans. [Learn more⁠(opens in a new window)](https://chatgpt.com/features/space/)

### Collaborative slides

Soon you’ll be able to create interactive slides with ChatGPT and your team, from a conversation or your own template. Multiple teammates and agents can edit the deck at the same time and leave comments. Present in ChatGPT or export to PowerPoint or Google Slides with formatting intact.

Available to all Pro, Business, and
OpenAI——
🟠 redditOpenAI launches GPT‑6.1 Sol and a $500/month ChatGPT Pro plan
OpenAI
ebytes11113
🟠 redditFirst impression about gpt6.1 sol
OpenAI
Individual_Art_516321
🟠 redditSol 6.1 Max and Blender for games — that just won't cut it.
OpenAI
Comfortable-Cat-96117781
🟠 redditGPT-6.1 Sol beat Pokemon Red's first gym in 244 turns, a new record on PokeBench, for $2.60
OpenAI
VibeCodyH15718
🟧 hnArtifical Analysis - GPT-6.1 Sol Replaces 6 Sol After 7 DaysFe2O320
🟧 hnGPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligencetheanonymousone8099
🟠 redditHow are we all finding 6.1 so far?
OpenAI
Chemical-Agency-39971214
🟠 reddit6.1 Sol - So far, pretty impressed
OpenAI
therealjerseytom17551
🟠 redditWhat was GPT-6-SOL?
OpenAI
Nortixon3819
🟠 redditWatched the DevDay stuff yesterday and for the first time I'm actually rethinking my $200 Claude Max sub.
ClaudeAI
OccasionNo4703011
🟠 redditGPT-6.1 Sol is now #1 on MathArena: 86.3% accuracy for $0.94, beating Astra’s 81.9% at $2.26
singularity
141_133721317
🟠 redditGPT 6.1 Sol vs GPT 6 Sol vs Opus 5.5 – Testing quota usage (Part 2)
ClaudeAI
Background-Web-631211
🟧 hnCafeBench: Sol 6.1 is great value, but not quite Opus levelzodwick50
🟠 redditGPT 6 Sol lasted 7 days before OpenAI replaced it with GPT 6.1 Sol
OpenAI
Top_Knee968710
🟠 redditsol 6.1 is actually pretty decent
OpenAI
Imapatato1210841
🟠 redditAs a Plus subscriber, Sol 6.1 is a game changer
OpenAI
penisbike6917559
🟠 redditPhoto-to-Blender benchmark: GPT-6 Astra won every photo, GPT-6.1 Sol scored 61 for 36 cents
OpenAI
smith200837
🟠 redditGPT-6.1 sol (max) scores 100% in frontier math 4
singularity
Southern-Break5505586109
🟠 redditFalling Stars – I didn't miss the code-x
ClaudeAI
Peleias03
🟠 redditDowngrading plans as long as I'm using 6.1 Sol?
OpenAI
DesiGrit1311
🟠 redditgpt 6.1 is the most annoying model i ever used.
OpenAI
Simple-Law5883013
🟠 reddit5.6 Sol is lazy AF
singularity
FlamaVadim06
🟠 redditInsane cost difference between Claude Opus 5.5 and Codex Sol 6.1
OpenAI
Electronic_Animal_555223
🟠 redditMicrosoft has accidentally revealed GPT-6.1 Sol uses the same base weights as GPT-6 Sol, but with only 2 inference passes instead of 3.
singularity
ResultBackground2450703108
🟠 reddit6.1-sol just disappeared from my codex
OpenAI
infinonsen24
🟠 redditTested OpenAI's "~50% faster" Codex update on my frozen benchmark. Didn't see it (yet)
ClaudeAI
Sufficient-Storage8711
🟠 reddit6x-Sol is so bad but Astra is too expensive
OpenAI
Late_Change5029033
🟠 redditOpenAI rolls out GPT-6 with Intelligent UI — ChatGPT answers can now include interactive tools
artificial
winer6662310
🟧 openaiAsana cuts model costs 76x in browser tests with GPT-6.1 Sol
Retrieved article excerpt

Open article · Retrieved 2026-10-09T19:36:30.956011+00:00

October 9, 2026

# Asana cuts model costs 76x in browser tests with GPT‑6.1 Sol

Using GPT‑6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.

[Contact sales](https://openai.com/contact-sales/)

White Asana logo over a blue layered-paper texture.

Company size: Enterprise

Region: North America

Industry: Technology

Products: Codex

76x

Lower estimated model costs with the optimized GPT-6.1 Sol workflow

5x

Faster browser runs with the optimized GPT-6.1 Sol workflow

$0.47

Average estimated model cost with the optimized GPT-6.1 Sol workflow

Loading…

Share

With GPT‑6 Astra in Codex running experiments, Asana optimized its browser agent’s workflow on GPT‑6.1 Sol to run 76x cheaper and 5x faster.

Asana helps customers automate work across business applications through [StackAI⁠(opens in a new window)](https://www.stackai.com/), a platform it [acquired⁠(opens in a new window)](https://asana.com/press/releases/pr/asana-acquires-stackai-adding-cross-system-execution-for-human-agent-teams/e7c73b97-ae8c-4e51-b927-189ccb184146). Using StackAI, customers can build workflows that navigate websites, fill out forms, and gather information without writing code. At Asana’s scale, small inefficiencies in these workflows add up.

Asana’s StackAI CTO, Frank Hidalgo, PhD, set out to make the browser agent faster and cheaper to run. He directed GPT‑6 Astra in Codex to investigate the agent, test improvements, and compare the results. Work he estimates would have taken one to two months by hand took about a week.

[Asana’s 144-run study⁠(opens in a new window)](https://asana.com/inside-asana/cut-browsers-agent-cost) tested GPT‑6.1 Sol and three other frontier models, called here Models A, B and C. The optimized workflow that emerged on GPT‑6.1 Sol averaged $0.47 in estimated model costs and about four minutes per run, 76x cheaper and 5x faster than the original production setup on Model B.

> “This is what teams of humans and agents look like in practice. An engineer set the direction, GPT-6 Astra ran the experiments, and the results went through Command to production. This demonstrates how Asana brings human and agent teams to life.”

—Arnab Bose, CPO at Asana

## Identifying browser-agent inefficiencies with GPT‑6 Astra

To move quickly, Hidalgo started by using GPT‑6 Astra in Codex to map the codebase and explain how the agent built each model request. GPT‑6 Astra discovered that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered, so every request resent that history at full price.

The agent also dropped older screenshots and trimmed text at nearly every step. Each edit altered the history, so caching the history alone would not have helped, and losing those facts could require the agent to revisit pages it had already read.

## From an estimated two months of research to one week with GPT‑6 Astra

Hidalgo reviewed GPT‑6 Astra’s proposed fixes and selected three to test:

- Extending caching to the agent’s browsing history
- Increasing the amount of text it could retain
- Removing screenshots in batches rather than at every step

GPT‑6 Astra began with quick tests to establish which variables mattered. Because the code was not designed for controlled experiments, it then refactored the code so one frontend and backend could support many workflows in parallel, each with its own settings.

Astra conducted the full study: history budgets of 120,000 and 480,000 characters and six caching and screenshot policies, each tested three times on each of the four models (see the table below). The best-performing policy allowed screenshots to accumulate to 20 before cutting back to the most recent one. This kept earlier history unchanged for longer stretches between removals. Combined with the larger history budget, it became the optimized workflow. Each configuration performed the same task: collecting six fields for each of 32 books from a public demo catalog, representative of what some Asana customers run in StackAI.

| **Model** | **Description** | **Price** |
| --- | --- | --- |
| **Model A** | A smaller, less expensive model from another frontier lab, released Fall 2025 | Half the price of GPT‑6.1 Sol |
| **Model B** | The model originally used in production, from the same lab as Model A, released Summer 2026 | Same price as GPT‑6.1 Sol |
| **Model C** | An updated version of Model B, released Fall 2026 | Same price as GPT‑6.1 Sol |
| **GPT‑6.1 Sol** | OpenAI’s model |  |

GPT‑6 Astra ran the workflows and examined the requests, usage records and outputs, and separate model sessions reviewed the work. Every session’s requests, data traces and results were recorded in [Command⁠(opens in a new window)](https://asana.com/product/command), Asana’s software delivery platform, so the team could review the complete study afterward. From Command, the findings were turned into tickets, then pull requests, and the changes went to production.

> “This would have taken me one to two months by hand. With GPT-6 Astra in Codex, it took about a week: I’d set a /goal before going to bed and review the results in the morning.”

—Frank Hidalgo, PhD, StackAI CTO at Asana

## Bringing model cost below $0.50 per run

For Model B, the optimization reduced estimated model cost from at least $36.21 (some original runs reached the step limit before they finished) to $1.24 per run, a reduction of 29x. The optimized workflow on GPT‑6.1 Sol was 2.6x cheaper still, at $0.47. Every run in the optimized workflow completed the task and returned the correct answer.

Means of 3 runs. ≥: the baseline includes capped runs, so its mean is a lower bound.

The two right folds compare against Model B optimized. Model B ran in phase 1, Model C and Sol 6.1 in phase 2 of the same study (dotted line).

On GPT‑6.1 Sol alone, with the larger history budget, the new caching and screenshot policy reduced cost 4x, from $1.97 to $0.47 per run. Each call was about 3x cheaper, because 89% of the input came from cache at 5% of the uncached price. Runs also became faster: at least 22.5 minutes on the original setup on Model B, roughly four minutes with the optimized workflow on GPT‑6.1 Sol.

Mean of 3 runs, SD whiskers. ≥: mean includes a capped or unfinished run, so the true value is at least this large.

Bars use the blue theme. Read caching effects against the 480k larger-budget bar.

Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.

Mean of 3 runs, SD whiskers. ≥: mean includes a capped or unfinished run, so the true value is at least this large.

Bars use the blue theme. Read caching effects against the 480k larger-budget bar.

Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.

The investigation also showed how history management affected whether the agent produced an answer at all. Giving GPT‑6.1 Sol more room to retain its browsing history increased the number of runs that produced an answer from three of 18 with the smaller history budget to all 18 with the larger budget, each with the correct answer. For Hidalgo, the business value is giving customers access to faster, more capable models while keeping operating costs sustainable.

> “Cost used to limit which models we could offer customers for these workloads. By making the agent more efficient, we can give customers a better, faster model while lowering our operating costs.”

—Frank Hidalgo, PhD, StackAI CTO at Asana

## Scaling experimentation and product testing

Asana has released the changes to browser navigation in StackAI and is developing tools to make similar experiments easier to repeat. Over time, the team plans to incorporate this testing into the platform’s evaluations, so customers and internal teams can compare cost, runtime, and answer quality when configuring their agents.

> “Shipping speed is no longer the bottleneck; human attention is. We’re close to a world where every engineer is a PM leading a fleet of agents.”

—Frank Hidalgo, PhD, StackAI CTO at Asana

Asana is now using GPT‑6 Astra in Codex to test product features before release: Astra navigates the platform, tries different inputs and reports bugs for human QA reviewers. Hidalgo views this as the foundation for a new software development lifecycle, with many cloud agent sessions testing features in parallel.

*The complete study is available on the* [*Asana*⁠(opens in a new window)](https://asana.com/inside-asana/cut-browsers-agent-cost) *and* [*StackAI*⁠(opens in a new window)](https://www.stackai.com/blog/how-stackai-by-asana-used-astra-codex-and-command-to-reduce-browser-agent-costs) *blogs.*

## Join the new era of work

More than 1 million businesses around the world are achieving meaningful results with OpenAI.

[Contact sales](https://openai.com/contact-sales/)

## Keep reading

Oracle — Cover

[How Oracle turns days of work into minutes with ChatGPT and Codex

Oct 8, 2026](https://openai.com/index/oracle/)

1Password > Card image > Fiber ridge close-up

[1Password increases engineering productivity 21% with Codex

Sep 8, 2026](https://openai.com/index/1password/)

loveholidays customer story art card v4

[How loveholidays is making everyone a builder with Codex

Aug 26, 2026](https://openai.com/index/loveholidays/)
OpenAI——

Interpretation history

Decision trace