2026-10-11 16:38 UTC

OpenAI claims Astra for Law's specialized search index and legal instructions raise legal-research correctness from 38.7% to 54.0% versus Astra with web search alone, providing a stronger foundation for legal workflows through restricted initial access, partner plugins, and a forthcoming API.

state: watchingheat: lowuncertainty: mediumknownscott: lowfrontier-models legal-agents rag enterprise-aiOpenAIHarveyLegoraFree Law Project
Surfaced 2026-09-19T08:26:48Z — priced heat=high at reprice: Strong cross-platform attention and renewed Reddit velocity now warrant high heat, even though the discussion still traces to the same launch rather than independent validation. This changes the attention judgment, not confidence in the benchmark or the offering's practical relevance to Scott.

What is this?

Astra for Law, launched 2026-09-17, is OpenAI's legal-industry configuration of its GPT-6 Astra frontier model: the model plus a specialized U.S. legal search index (230M+ URLs of case law, statutes and regulations, reported as drawing on Free Law Project/CourtListener data) plus legal-analysis instructions. It rolled out first through restricted 'Trusted Access' to selected U.S. firms, with an API for legal-tech partners such as Harvey and Legora still listed as forthcoming. OpenAI reports the configuration lifts overall correctness on Vals AI's 200-question Legal Research Bench from 38.7% (same model, plain web search) to 54.0% at matched reasoning effort; the supplied secondary coverage uniformly repeats these as OpenAI's own numbers, with none independently validating the claim or attributing the gain to the index versus the instructions. The snippets do not cover the case's two later first-party items — the Ironclad computer-use collaboration and the newly added launch-page head-to-head vs Claude Fable 5.1 — so the head-to-head's provenance (post-launch edit vs. earlier unrendered capture) remains undated on this evidence.

Why it matters to Scott

Still held territory: the launch's own ablation (same GPT-6 Astra model, +15.3 points from a curated legal index plus instructions) re-evidences his capability-symmetry and RAG/wiki substrate-rule positions that context architecture and the compiled corpus — not raw model capability — decide domain correctness, while the restricted-access, lawyer-retains-judgment deployment repeats the cognitive-exoskeleton division of labour already recorded at grounding. The new Claude Fable head-to-head is a textbook instance of the stale-authority claim his witness-not-oracle contract tests for, but all three performance claims remain OpenAI-reported — seller-class under his evidence-class ladder, with no independent validation and no inspectable API — so nothing here challenges, extends, or makes actionable what he builds or argues.
ip:concept.capability-symmetryip:framework.rag-wiki-substrate-ruleip:framework.cognitive-exoskeleton-patternip:source.witness-not-oracle-ebookip:concept.evidence-class-ladderip:concept.capability-auditradar:openai-chatgpt-financial-servicesradar:thomson-reuters-frontier-modelradar:harvey-rlm-ma-diligenceradar:perplexity-numeric-citation-auditradar:concept.ragradar:concept.enterprise-airadar:concept.computer-use-agentsradar:concept.agentic-rlradar:concept.citation-verificationradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • domain retrieval index vs general web search accuracy RAG
  • vertical model configuration enterprise wrapper moat product pattern
  • lab self-reported benchmarks independent validation receipts
  • cognitive exoskeleton professional judgment automation boundary
  • computer-use agent RL training on real workflow tasks
  • citation integrity grounded holdings agent memory wiki

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 592h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 00:00⭐ origin directly observedIntroducing Astra for Law
OpenAI on openai
—
09-17 20:17first on hacker news · published · +20.3hAstra for Law
vertigoruntime
—
09-18 09:40first on r/singularity · published · +33.7hOpenAI: "Introducing GPT-6 Astra for Law" (new model "gpt-6-astra-law")
borowcy
—
10-06 10:00first on openai · published · +466.0hAdvancing computer use with Ironclad
OpenAI
—
09-17 20:17amplified on hacker news 👑hn.story.49745940
vertigoruntime
peak 589 · 689 comments · 71% of case engagement
09-18 09:40amplified on r/singularityreddit.post.1wjlmzx
borowcy
peak 736 · 207 comments · 29% of case engagement
10-07 00:44amplified on hacker newshn.story.49986385
apetresc
peak 1 · 0 comments · 0% of case engagement
09-17 20:21our radar first saw it · +20.4hdiscovery anchor: hn.story.49745940—
09-19 08:26reached heat=high · +56.5h · via ledger——
pace: p95 vs 1032 stories at the 336h mark (now 592h old) — ahead of ai-stupid-level-benchmark-drift (1.0x), behind openai-chatgpt-weekly-prompt-caps (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAstra for Law
Retrieved article excerpt

Open article · Retrieved 2026-09-17T20:23:07.834326+00:00

September 17, 2026

[Company](https://openai.com/news/company-announcements/)

# Introducing Astra for Law

Our most powerful model, configured into a new AI foundation for law.

[Explore solutions for law](https://openai.com/solutions/industries/law/)[Contact Legal sales](https://openai.com/business/contact-sales-legal/)

Loading…

Share

Today, we’re introducing Astra for Law: a new foundation for law firms and legal technology companies to build AI products and workflows around their expertise. It combines GPT‑6 Astra, our latest and most powerful model, with settings, tools, and context tailored for professional legal work.

API customers including Harvey and Legora will be able to build on Astra for Law, bringing this intelligence into their own products and workflows. As our frontier models advance, we’ll bring these legal capabilities to our latest models.

We are also expanding our work on privacy and governance to give law firms specific controls for confidential client work. Firms can also customize Astra for Law using our 26 new ecosystem plugins that connect ChatGPT to the specialist tools firms already use, like Relativity and Clio.

## Frontier intelligence for law

Astra for Law combines GPT‑6 Astra with a powerful legal search index and instructions for legal analysis and writing. Together, they amplify Astra’s capabilities across the legal practice, while giving firms and legal technology companies the freedom to build their own applications and workflows.

### Legal research: from facts to a supported answer

Our new legal search index is one of the tools Astra for Law can use. Legal research often begins with finding the exact right authority, locating the relevant passages, and understanding how relevant and binding they are to the situation at hand. The index helps Astra for Law do that work, and complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters.

By using the legal search index, Astra for Law can search U.S. case law, statutes, regulations, court rules, and administrative decisions across a corpus of more than 230 million URLs, with sources added daily. Our work with Free Law Project, the nonprofit behind CourtListener, brings its case-law collection covering [more than 99.9% of published U.S. precedential case law⁠(opens in a new window)](https://wiki.free.law/c/courtlistener/help/data-coverage/case-law) into this research experience.

To measure how this configuration improves legal research, we tested Astra for Law’s complete setup on 200 U.S. legal research questions from the private validation set of [Vals AI’s Legal Research Bench⁠(opens in a new window)](https://vals.ai). This benchmark measures how well the model can find relevant sources and passages, and how well its research answers meet the evaluation criteria.

At the highest reasoning effort for both systems, Astra for Law passed the evaluation’s overall correctness check on 54.0% of questions, compared with 38.7% for GPT‑6 Astra using web search alone – a 40% relative improvement. Astra for Law also produces more comprehensive answers.

On case-law-focused questions, Astra for Law found 24% more reference cases than GPT‑6 Astra using web search alone at the highest reasoning effort. On the audited set of target passages, it retrieved up to 54% more relevant passages from the correct court opinions, when comparing the systems at the same reasoning effort.

The result is a stronger research foundation for advising on a deal, assessing a dispute, or developing a legal strategy, with reliable authorities the lawyer can examine for herself.

Astra for Law and GPT‑6 Astra’s performance on the Vals AI Legal Research Bench validation set, across reasoning effort settings.

### Improving performance on end-to-end legal workflows

Legal research is only the first step. Custom instructions for legal analysis and writing guide Astra for Law in applying that research to the client’s facts, developing arguments or deal terms, and identifying weaknesses and uncertainty. That can mean distinguishing a court’s holding from its other observations, addressing cases that weaken an argument, or explaining how a contract exception shifts risk between the parties.

For example, when prompted to identify good law with similar fact patterns, Astra for Law could both pinpoint relevant precedent and match fact patterns better than other frontier models:

Astra for Law will be initially offered to selected law firms through Trusted Access in ChatGPT and Codex, and will be coming soon to the API. It will appear in the model picker as “GPT‑6 Astra Law” and in the API as gpt-6-astra-law.

1 of 3

> “We were grateful to preview early versions of Astra for Law, which were built with legal use-cases in mind. Across both litigation and transactional matters, the models demonstrated impressive research depth and sensitivity to authority. Even at this early stage, they felt like a significant step toward legal-focused AI that is carefully grounded in research that is both current and comprehensive.”

John Savva, Partner, **Sullivan & Cromwell**

> “In our early testing, Astra for Law showed strength across key aspects of legal research: grounding answers in on-point authorities, citing with precision, and offering practical, advisory guidance.”

Niko Grupen, Head of Applied Research at Harvey

> “Our clients rely on us for judgment on the deals, disputes, and regulatory challenges that define their industries. We see an extraordinary opportunity in partnering with OpenAI on how its technology can enhance and elevate our capabilities. Together, we are shaping how frontier AI applies to the most demanding legal work, starting with transactional due diligence. Our goal is to put more of our lawyers' insight to work earlier in the process, so clients get faster execution, sharper analysis, and the industry-leading counsel they depend upon.”

**Ropes & Gray**

- Sullivan & Cromwell
- Harvey
- Ropes & Gray

- Sullivan & Cromwell
- Harvey
- Ropes & Gray

## Legal-grade trust and controls

Law firms need to protect client confidences and control how AI is used in their practice. We’ve created a special Trusted Access Program for eligible law firms to give lawyers and people working under their supervision access to Astra for Law for professional legal work. For eligible firms, the offering includes Zero Data Retention (ZDR) on our API, and usage of ChatGPT Enterprise is excluded from human review by default.

We are also working with **Latham & Watkins**, a leader in AI governance, to design for information permissions, ethical walls, client instructions, and firm oversight.

> “As AI becomes more capable, so too does the ability to deploy it in environments that demand rigorous governance, oversight, and accountability. This collaboration builds on **Latham’s** broader investments in AI development, governance, and infrastructure across the firm, which underpin our enterprise-wide strategy for responsible AI.”

Michael Rubin, Chair of **Latham’s** AI Strategy Committee

## Build ChatGPT around your firm’s expertise

With frontier intelligence and the right controls, firms can turn their own precedents, methods, and judgment into AI tools and workflows built to their standards.

Working with selected firms, our forward-deployed engineers have been adapting ChatGPT Enterprise with custom interfaces and integrations to proprietary data, creating tools for each firm’s workflows:

- **Sullivan & Cromwell** built an agreement analyzer that brings the firm’s negotiating playbooks and selected precedents into the review of a new deal. It helps lawyers spot risks that emerge when provisions are read together, then turns those findings into proposed redlines and draft client advice they can challenge and refine.
- **Ropes & Gray** built a deal diligence system around how its lawyers work through a data room and decide what matters to the deal. It helps them trace findings back to the source and pinpoint questions that could affect an acquisition, such as whether key customer contracts require notice or consent.
- **Cooley** built GO Public to bring its capital markets expertise into how companies prepare to go public, from drafting the IPO filing to identifying the risks that deserve management’s attention. When the deal changes, it carries that change across the filing so lawyers can review the implications together.

1 of 3

> “Our collaboration with OpenAI has allowed us to rethink how this work gets done – moving lawyers and management teams more quickly through intensive preparation and into questions that require judgment, market experience and strategic thinking.”

Dave Peinsipp, partner and co-chair of **Cooley’s** global capital markets group

> “AI is fundamentally transforming the way legal services are delivered, and **Sullivan & Cromwell** is committed to being at the forefront of that evolution. Our team is developing custom AI applications aligned with the way our lawyers work. OpenAI's world-class engineering expertise is helping us bring our vision to life through an initial test application. Our collaboration refined the application, enhanced its performance and prepared it for broad deployment. The experience demonstrated the value of purpose-built AI tools supported by a strong technology platform and shaped by our firm's standards. We are excited to continue exploring how such tools can enhance the way our lawyers work and help us deliver even greater value to our clients.”

Robert Giuffra Jr. and Scott Miller, Co-Chairs, **Sullivan & Cromwell**

> “Our clients rely on us for judgment on the deals, disputes, and regulatory challenges that define their industries. We see an extraordinary opportunity in partnering with OpenAI on how its technology can enhance and elevate our capabilities. Together, we are shaping how frontier AI applies to the most demanding legal work, starting with transactional due diligence. Our goal is to put more of our lawyers' insight to work earlier in the process, so clients get faster execution, sharper analysis, and the industry-leading counsel they depend upon.”

**Ropes & Gray**

- Cooley
- Sullivan & Cromwell
- Ropes & Gray

- Cooley
- Sullivan & Cromwell
- Ropes & Gray

Firms can build with their own teams and partner products, with permitted sources and review processes defined for the work.

## Frontier intelligence, connected to the most trusted tools in legal

We’re proud to work with the specialist companies who are advancing legal AI. Today, we’re launching [26 partner-built plugins⁠](https://openai.com/business/plugins/?tab=plugins-legal) that help firms go deeper with the tools and knowledge they already use. These plugins cover the practice and business of law. With iManage, a lawyer can draft a negotiation brief in ChatGPT and save it to the matter file; Intapp can surface activities that may need a time entry for review; DeepJudge can bring prior deals into a comparison. Thomson Reuters is bringing HighQ matter context into ChatGPT and previewing a forthcoming CoCounsel Legal connector.

> “As AI becomes more open and interoperable, the value is not in connectivity alone. Legal professionals need more than access to information. They need trusted intelligence, relevant enterprise and matter context, purpose built legal capabilities, and the governance required for high stakes work. That is what Thomson Reuters delivers through HighQ and CoCounsel Legal. CoCounsel remains the trusted professional AI system designed to help complete that work. Our work with OpenAI helps make these capabilities available in the environments customers choose, while preserving the accuracy, confidentiality, and accountability they depend on.”

Joel Hron, Chief Technology Officer, Thomson Reuters

The launch includes 9 community plugins from lawyers and legal engineers at LegalQuants, LECG, and Skills.law, with 47 custom skills that pract
vertigoruntime589689
🟧 openai ⭐Introducing Astra for Law
Retrieved article excerpt

Open article · Retrieved 2026-10-07T06:21:50.918770+00:00

September 17, 2026

[Company](https://openai.com/news/company-announcements/)

# Introducing Astra for Law

Our most powerful model, configured into a new AI foundation for law.

[Explore solutions for law](https://openai.com/solutions/industries/law/)[Contact Legal sales](https://openai.com/business/contact-sales-legal/)

Loading…

Share

Today, we’re introducing Astra for Law: a new foundation for law firms and legal technology companies to build AI products and workflows around their expertise. It combines GPT‑6 Astra, our latest and most powerful model, with settings, tools, and context tailored for professional legal work.

API customers including Harvey and Legora will be able to build on Astra for Law, bringing this intelligence into their own products and workflows. As our frontier models advance, we’ll bring these legal capabilities to our latest models.

We are also expanding our work on privacy and governance to give law firms specific controls for confidential client work. Firms can also customize Astra for Law using our 26 new ecosystem plugins that connect ChatGPT to the specialist tools firms already use, like Relativity and Clio.

## Frontier intelligence for law

Astra for Law combines GPT‑6 Astra with a powerful legal search index and instructions for legal analysis and writing. Together, they amplify Astra’s capabilities across the legal practice, while giving firms and legal technology companies the freedom to build their own applications and workflows.

### Legal research: from facts to a supported answer

Our new legal search index is one of the tools Astra for Law can use. Legal research often begins with finding the exact right authority, locating the relevant passages, and understanding how relevant and binding they are to the situation at hand. The index helps Astra for Law do that work, and complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters.

By using the legal search index, Astra for Law can search U.S. case law, statutes, regulations, court rules, and administrative decisions across a corpus of more than 230 million URLs, with sources added daily. Our work with Free Law Project, the nonprofit behind CourtListener, brings its case-law collection covering [more than 99.9% of published U.S. precedential case law⁠(opens in a new window)](https://wiki.free.law/c/courtlistener/help/data-coverage/case-law) into this research experience.

To measure how this configuration improves legal research, we tested Astra for Law’s complete setup on 200 U.S. legal research questions from the private validation set of [Vals AI’s Legal Research Bench⁠(opens in a new window)](https://vals.ai). This benchmark measures how well the model can find relevant sources and passages, and how well its research answers meet the evaluation criteria.

At the highest reasoning effort for both systems, Astra for Law passed the evaluation’s overall correctness check on 54.0% of questions, compared with 38.7% for GPT‑6 Astra using web search alone – a 40% relative improvement. Astra for Law also produces more comprehensive answers.

On case-law-focused questions, Astra for Law found 24% more reference cases than GPT‑6 Astra using web search alone at the highest reasoning effort. On the audited set of target passages, it retrieved up to 54% more relevant passages from the correct court opinions, when comparing the systems at the same reasoning effort.

The result is a stronger research foundation for advising on a deal, assessing a dispute, or developing a legal strategy, with reliable authorities the lawyer can examine for herself.

Astra for Law and GPT‑6 Astra’s performance on the Vals AI Legal Research Bench validation set, across reasoning effort settings.

### Improving performance on end-to-end legal workflows

Legal research is only the first step. Custom instructions for legal analysis and writing guide Astra for Law in applying that research to the client’s facts, developing arguments or deal terms, and identifying weaknesses and uncertainty. That can mean distinguishing a court’s holding from its other observations, addressing cases that weaken an argument, or explaining how a contract exception shifts risk between the parties.

For example, when prompted to identify good law with similar fact patterns, Astra for Law could both pinpoint relevant precedent and match fact patterns better than other frontier models:

Prompt

Our client signed a five-year agreement to manufacture a retailer’s private-label products after being told the retailer’s comparable suppliers historically averaged about $8 million in annual orders. The contract commits the retailer to $2 million in annual purchases and leaves additional orders to its discretion. Before signing, our client knew the retailer was revising how it allocated orders among suppliers. Orders have barely exceeded the minimum. Internal records suggest the historical average was much lower. The retailer says our client accepted both the minimum and the new process. Find the closest factual precedent and write a memo on our client’s misrepresentation claim.

##### Claude Fable 5.1

MEMORANDUM — Misrepresentation Claim Against Retailer

Closest precedent. Paragon Metals Holdings v. Smith, C.A. No. N21C-12-090-SKR-CCLD (Del. Super. Aug. 13, 2025). A buyer alleged the seller concealed that two major customers would materially cut volumes. After a bench trial, the court rejected the fraud claim, concluding the buyers failed to conduct sufficient diligence and ignored repeated, if inartful, disclosures of the very information underlying their claim. Notably, the court assumed certain representations may have been false yet still ruled for the sellers because the buyers could not prove justifiable reliance.

Application. Our strongest fact is that the "$8 million average" was a statement of historical fact, not a forecast, and internal records apparently show it was false. That distinguishes the ordinary "predictions aren't actionable" defense. The $2 million minimum does not disclaim the historical figure; it allocates future risk, not past fact.

Our weakest fact mirrors Paragon: the client knew allocation methodology was changing and signed anyway. The retailer will argue that knowledge made reliance on old averages unreasonable and that the minimum captured the bargained-for risk. Anti-reliance or integration clauses in the agreement would compound this.

Recommendation. Plead fraudulent inducement (and negligent misrepresentation alternatively) centered on the false historical average; develop evidence that the process change was presented as neutral to volume. Confirm governing law—Delaware is notably strict on reliance.

##### Astra for Law

Our client has a viable claim based on false history, not a right to $8 million in annual orders.

The closest analogue is Oliver Wyman, Inc. v. Eielson, 282 F. Supp. 3d 684, 700–03 (S.D.N.Y. 2017), applying Massachusetts law. Consulting-business sellers accepted four-year employment agreements guaranteeing $425,000, with discretionary increases, after representations of approximately $1 million average partner compensation. They knew the compensation system was changing. The court denied summary judgment on fraud and negligent-misrepresentation counterclaims concerning the historical average: evidence disputed its accuracy, and the contractual floor did not contradict past results.

Similarly, the retailer’s approximate $8 million historical average is a verifiable factual assertion. Accepting a $2 million minimum and discretionary additional orders does not necessarily surrender reliance on truthful historical information. But Oliver Wyman emphasized that no evidence showed the revised system would fundamentally change compensation or render historical averages irrelevant. Id. at 701–02. The retailer’s strongest defense is proof our client understood the allocation changes would make past supplier volumes uninformative—not simply that changes were announced.

For fraud, our client must prove knowing material falsity, intent to induce signing, actual reasonable reliance, and resulting loss. Negligent misrepresentation may instead rest on deficient care in supplying business information. Obtain supplier-level records matching the represented period and comparator group, the speaker’s contemporaneous communications, and our client’s decision records.

Given the same prompt, Astra for Law returned two closely matching precedents; in the litigation example, Claude Fable 5.1 returned a holding that had been reversed on appeal, while in the transactional example it reported finding no such case.

Astra for Law will be initially offered to selected law firms through Trusted Access in ChatGPT and Codex, and will be coming soon to the API. It will appear in the model picker as “GPT‑6 Astra Law” and in the API as gpt-6-astra-law.

> “We were grateful to preview early versions of Astra for Law, which were built with legal use-cases in mind. Across both litigation and transactional matters, the models demonstrated impressive research depth and sensitivity to authority. Even at this early stage, they felt like a significant step toward legal-focused AI that is carefully grounded in research that is both current and comprehensive.”

John Savva, Partner, **Sullivan & Cromwell**

> “In our early testing, Astra for Law showed strength across key aspects of legal research: grounding answers in on-point authorities, citing with precision, and offering practical, advisory guidance.”

Niko Grupen, Head of Applied Research at Harvey

1 of 2

> “We were grateful to preview early versions of Astra for Law, which were built with legal use-cases in mind. Across both litigation and transactional matters, the models demonstrated impressive research depth and sensitivity to authority. Even at this early stage, they felt like a significant step toward legal-focused AI that is carefully grounded in research that is both current and comprehensive.”

John Savva, Partner, **Sullivan & Cromwell**

> “In our early testing, Astra for Law showed strength across key aspects of legal research: grounding answers in on-point authorities, citing with precision, and offering practical, advisory guidance.”

Niko Grupen, Head of Applied Research at Harvey

- Sullivan & Cromwell
- Harvey

- Sullivan & Cromwell
- Harvey

## Legal-grade trust and controls

Law firms need to protect client confidences and control how AI is used in their practice. We’ve created a special Trusted Access Program for eligible law firms to give lawyers and people working under their supervision access to Astra for Law for professional legal work. For eligible firms, the offering includes Zero Data Retention (ZDR) on our API, and usage of ChatGPT Enterprise is excluded from human review by default.

We are also working with **Latham & Watkins**, a leader in AI governance, to design for information permissions, ethical walls, client instructions, and firm oversight.

> “As AI becomes more capable, so too does the ability to deploy it in environments that demand rigorous governance, oversight, and accountability. This collaboration builds on **Latham’s** broader investments in AI development, governance, and infrastructure across the firm, which underpin our enterprise-wide strategy for responsible AI.”

Michael Rubin, Chair of **Latham’s** AI Strategy Committee

## Build ChatGPT around your firm’s expertise

With frontier intelligence and the right controls, firms can turn their own precedents, methods, and judgment into AI tools and workflows built to their standards.

Working with selected firms, our forward-deployed engineers have been adapting ChatGPT Enterprise with custom interfaces and integrations to proprietary data, creating tools for each firm’s workflows:

- **Sullivan & Cromwell** built an agreement analyzer that brings the firm’s negotiating playbooks and selected precedents into the revie
OpenAI——
🟠 redditOpenAI: "Introducing GPT-6 Astra for Law" (new model "gpt-6-astra-law")
singularity
borowcy732204
🟧 openaiAdvancing computer use with Ironclad
Retrieved article excerpt

Open article · Retrieved 2026-10-06T23:22:39.610153+00:00

October 6, 2026

[Company](https://openai.com/news/company-announcements/)[Publication](https://openai.com/research/index/publication/)

# Advancing computer use with Ironclad

How a research collaboration in contracting is helping us train and evaluate AI agents on complex professional work.

Loading…

Share

When we introduced [GPT‑6 Astra](https://openai.com/index/gpt-6-astra/), we demonstrated how far our models have come in using computers for professional work, from preparing documents to testing websites. Our next goal is to make agents more capable and efficient at using specialized software to solve complex business problems. We’re exploring how to train models to understand a company’s business rules, execute multi-step workflows, and verify that their work meets the original requirements.

To accelerate this research, we’re partnering directly with a small number of software companies that understand these workflows best. Together, we’re identifying challenging, high-value tasks and turning them into research problems for training and evaluating our models, ultimately making our models more capable and useful in real-world business applications.

Our first partner is Ironclad, a leader in AI contracting. Working closely with Ironclad’s team, we’ve developed tasks that require agents to configure agreements, approvals, and reusable legal terms, and demonstrated progress on these complex workflows. Ironclad’s expertise has been instrumental in defining what success looks like and bringing real customer needs directly into frontier model development. We’re grateful for their partnership and excited to share what we’ve accomplished together.

GPT‑6 Astra is our first frontier model trained on Ironclad tasks. On our research evaluation, its average score was 32% higher than GPT‑5.6 Sol’s, while estimated time per attempt was 48% lower.[1](https://openai.com/index/advancing-computer-use-with-ironclad/#citation-bottom-1)

## Turning contracting workflows into training tasks

Consider a legal operations team setting up a process for buying software. Finance may need to approve purchases above a certain amount, Security may need to review certain requests, and Legal may need to review nonstandard terms. The person setting up the process has to turn that short list into an intake form, document templates, approval rules, and a record of the final agreement.

An AI agent doing the same work has to keep those requirements in view as it moves through the software. For example, it must configure Finance approval above the spending threshold and check that requests above and below it follow the right paths. Getting individual steps right is not enough: the finished process must work across the situations it was designed to handle.

Ironclad employees and people who use Ironclad at OpenAI helped our researchers identify 11 tasks across legal, commercial, and procurement work. These included tasks like setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause so that it reflects the jurisdiction a requester selects. We estimate that this work would take an experienced user about 30 to 40 minutes per task, on average.

We evaluated each task against 8 to 50 criteria, depending on its complexity. This let us see which parts a model got right and where it fell short. Ironclad also provided hosted software environments of their product where the models could practice these tasks. Our researchers developed synthetic training tasks[2](https://openai.com/index/advancing-computer-use-with-ironclad/#citation-bottom-2) around representative workflows and used reinforcement learning to help the models improve through practice and feedback. This combination—tasks selected with people who know the work, detailed criteria, and a place for models to practice—ensures we are improving on real-world tasks that are most important to our customers.

## How Astra performed

We compared Astra and GPT‑5.6 Sol using Max reasoning for Astra and High reasoning for Sol, the settings where each model scored highest. Across the 11 research tasks, Astra’s average score was 55.0%, compared with 41.6% for GPT‑5.6 Sol, while estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra.[3](https://openai.com/index/advancing-computer-use-with-ironclad/#citation-bottom-3) An internal model used in the development of Astra achieved an even stronger 63.7% on these tasks, and we aim to bring these further gains to future models.

|  |  |  |
| --- | --- | --- |
| **Metric** | **GPT‑5.6 Sol** High reasoning | **GPT‑6 Astra** Max reasoning |
| Mean rubric score | 41.6% | 55.0% |
| Estimated average time per attempt | 37.0 minutes | 19.2 minutes |

Astra met more of the task requirements with less simulated time per attempt. The two clips below show how the models handled the same research task. Astra (first) met about 94% of the task criteria in an estimated 20 minutes; GPT‑5.6 Sol (second) met about 85% in an estimated 32 minutes.

**GPT‑6 Astra** met about 94% of the task criteria in an estimated 20 minutes.

**GPT‑5.6 Sol** met about 85% of the task criteria in an estimated 32 minutes.

## What this collaboration means for Ironclad

Ironclad’s role in this work reflects a challenge its customers know well: a contracting process must handle exceptions while preserving the rules a business depends on. If an agent loses track of one of those rules halfway through a task, that limits what a software company can confidently ask it to do. This underscores why human oversight still matters as agents get better at complex contracting tasks, and why a full contracting platform remains essential. Working with our researchers gives Ironclad a way to bring these problems into the development of the underlying model, where we can study them together.

As models become more capable, Ironclad has an opportunity to use them in more of its own products. For legal and business teams, that could mean AI taking on more of the effort involved in complex contracting while preserving the controls those teams rely on.

> “For AI to be genuinely useful in complex areas of contracting, it has to do more than complete individual actions. Agents need to understand the full contracting lifecycle, including how business workflows connect while preserving the controls teams rely on. Our collaboration with OpenAI brings these real-world challenges into the research process and helps move the technology forward.”

—Sunita Verma, Chief Technology Officer, Ironclad

## Work with us to advance AI for complex professional work

We’re inviting a small number of software companies to work directly with our research and engineering teams on important professional tasks that today’s agents still can’t reliably complete. If your team has one, we’d like to see a concrete example: what you’re asking the agent to do, evidence of where it fails, and how you would judge a successful result. Partners should also be able to bring people who know the work deeply, a secure environment for testing, and data that can be safely used for research. This gives us a starting point for investigating the failure and measuring whether the model improves.

If that describes your team,[**apply to explore a research collaboration**](https://openai.com/form/computer-use-research-collaboration/).

- [2026](https://openai.com/news/?tags=2026)

## Author

OpenAI

## Footnotes

1. 1

   These results cover the 11 research tasks, not all Ironclad workflows. The times are simulated estimates based on assumed model processing and generation speeds, not measured customer time savings.
2. 2

   We created simulated tasks from contracts publicly available in the SEC’s EDGAR database after applying filters designed to remove personal information. We did not use OpenAI customer data, OpenAI’s internal contracts, or nonpublic Ironclad customer data or contracts for training or evaluation.
3. 3

   See the research-evaluation footnote above.

## Keep reading

[View all](https://openai.com/news/)

Sharing AI progress in mathematics — Cover

[Sharing AI progress in mathematics

ResearchOct 6, 2026](https://openai.com/index/sharing-ai-progress-in-mathematics/)

Atlassian partnership card — official white SVGs on Neutral 088

[Atlassian and OpenAI expand their partnership

CompanyOct 6, 2026](https://openai.com/index/atlassian-partnership/)

Albertsons | How Albertsons Companies is reimagining retail from the inside out | Cover

[How Albertsons Companies is reimagining retail from the inside out

CompanyOct 1, 2026](https://openai.com/index/albertsons-reimagining-retail/)
OpenAI——
🟧 hnAdvancing Computer Use with Ironcladapetresc10

Interpretation history

Decision trace