🔍 Read the full analysis: OpenAI Trains Agents Inside Software. Take Time To Read Ironclad’s Terms on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI’s October 6 post describes training GPT-6 Astra on legal, commercial and procurement workflows in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while its estimated completion times were simulations, not measured customer productivity gains.
OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The results offer an early test of agents working through complex business workflows, but Astra met an average 55% of task criteria, and the reported time savings are simulated rather than measured in customer use.
OpenAI said Ironclad staff and OpenAI employees selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and adapting reusable contract clauses to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task. Depending on complexity, each task was judged against 8 to 50 criteria; the score measures the share of criteria met, not the share of tasks completed successfully.
For the reported comparison, GPT-6 Astra scored 55.0% on the average share of criteria met, against 41.6% for GPT-5.6 Sol in high-effort mode. OpenAI also reported estimated times per attempt of 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%, while Astra met about 94% of criteria on one showcase task. Those figures describe a limited set of research tasks, not general performance across Ironclad or other software.
OpenAI said Ironclad provided hosted product copies for model practice. It said training tasks were synthesized from contracts publicly filed in the SEC’s EDGAR database and filtered to remove personal information. OpenAI stated it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. The post also says the time figures are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The research addresses a practical hurdle for workplace agents: completing several linked steps while retaining a company’s rules and checking that the final result satisfies them. In contracting and procurement, a missed condition can matter more than a middling overall score. If an approval process omits a required Finance, Security or Legal review, the workflow may fail at the point it is meant to control.
OpenAI’s reported 55% average criteria score is therefore not evidence that the model can independently handle these tasks reliably. The result suggests progress on a difficult benchmark, but the criteria-based scoring and limited sample leave room for important errors. OpenAI’s own post says human oversight remains necessary when an agent may lose track of a business rule.
The work could also affect software companies beyond Ironclad. OpenAI is asking a small number of vendors to contribute hard examples, knowledgeable staff, secure test environments and research-usable data. If agents become more capable inside specialized products, vendors may gain a new way to make their services useful. They may also face a shift in how customers interact with them: people could issue instructions to an agent instead of using the product’s screens. The vendor’s underlying rules, records, controls and audit trail would then carry more weight than its interface.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Built
The October 6 post described a research setup rather than a general product launch. OpenAI and Ironclad chose tasks that reflect work performed in contract-management software, then assessed model outputs against task-specific rubrics. The test covered 11 selected workflows; it does not establish how often agents can complete the full range of work that Ironclad customers perform.
OpenAI described the model training as practice in hosted copies of Ironclad’s product, using synthetic tasks built from public SEC-filed contracts. The company said it did not use non-public Ironclad customer data. The distinction matters because the work involves sensitive business processes, and the post’s account of the data sources is a statement from OpenAI, not an independent audit described in the source material.
The post’s final section invites other software companies to work with OpenAI on tasks current agents cannot reliably complete. It asks prospective partners to bring concrete failure cases, subject-matter expertise, secure testing environments and data that can safely be used for research. The stated aim is to train and evaluate agents against real business rules and multi-step workflows, rather than only generic computer-use tasks.
contract management software with AI integration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The post does not establish how Astra would perform across the full range of Ironclad workflows, on live customer accounts or under production conditions. The test involved 11 research tasks, and the average criteria score does not identify which particular requirements were missed on each task. The source material also does not provide a breakdown of error severity or independent verification of the results.
It remains unclear whether the estimated time figures correspond to any real customer productivity gains. OpenAI explicitly characterized them as simulations, and the comparison does not establish whether an agent plus human review would be faster or less costly than an experienced person completing and checking the work. The source also does not describe a customer deployment, release timeline or commercial arrangement between the companies.
OpenAI’s statements about training data and the limits of customer data use are attributable to the company. The source material does not mention an independent audit of the data or test environment. Nor does it specify which other software companies may participate in the proposed partner program.
AI-powered procurement workflow tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What OpenAI and Vendors May Test
OpenAI has invited a small number of software companies to bring difficult workflows to future research partnerships. The next developments to watch are whether other vendors join, what tasks they choose, how their results are scored and whether OpenAI publishes more detail on failure types and verification methods.
For any move from research to customer use, buyers will need clearer information about which rules an agent can handle, what happens when it misses a requirement, and who checks its work. The October 6 post describes training and evaluation work; it does not announce that customers can hand over contract workflows to Astra without supervision.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad announce?
OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The post also invited other software companies to discuss similar research partnerships.
Does the 55% result mean Astra completed 55% of tasks?
No. OpenAI reported that Astra met an average 55% of the criteria in the task rubrics. That is not the percentage of tasks completed, and it does not show that each workflow was safe or ready for use without review.
Did Astra cut customer task times by nearly half?
The post does not report measured customer savings. OpenAI’s figures of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol are simulated estimates based on assumed processing and generation speeds.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. It stated that it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can companies use Astra to manage contracts now?
The source describes research and testing, not a general customer deployment. OpenAI’s post says human oversight remains necessary, and it does not provide a release timeline for unsupervised contract or procurement work.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
