AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Trains Agents Inside Software. Take Time To Read Ironclad’s Terms on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s October 6 post describes training GPT-6 Astra on legal, commercial and procurement workflows in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while its estimated completion times were simulations, not measured customer productivity gains.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The results offer an early test of agents working through complex business workflows, but Astra met an average 55% of task criteria, and the reported time savings are simulated rather than measured in customer use.

OpenAI said Ironclad staff and OpenAI employees selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and adapting reusable contract clauses to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task. Depending on complexity, each task was judged against 8 to 50 criteria; the score measures the share of criteria met, not the share of tasks completed successfully.

For the reported comparison, GPT-6 Astra scored 55.0% on the average share of criteria met, against 41.6% for GPT-5.6 Sol in high-effort mode. OpenAI also reported estimated times per attempt of 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%, while Astra met about 94% of criteria on one showcase task. Those figures describe a limited set of research tasks, not general performance across Ironclad or other software.

OpenAI said Ironclad provided hosted product copies for model practice. It said training tasks were synthesized from contracts publicly filed in the SEC’s EDGAR database and filtered to remove personal information. OpenAI stated it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. The post also says the time figures are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers.

At a glance
reportWhen: Published October 6; research results a…
The developmentOpenAI described training a frontier model on contract-management workflows inside Ironclad software and invited other software companies to partner on similar research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The research addresses a practical hurdle for workplace agents: completing several linked steps while retaining a company’s rules and checking that the final result satisfies them. In contracting and procurement, a missed condition can matter more than a middling overall score. If an approval process omits a required Finance, Security or Legal review, the workflow may fail at the point it is meant to control.

OpenAI’s reported 55% average criteria score is therefore not evidence that the model can independently handle these tasks reliably. The result suggests progress on a difficult benchmark, but the criteria-based scoring and limited sample leave room for important errors. OpenAI’s own post says human oversight remains necessary when an agent may lose track of a business rule.

The work could also affect software companies beyond Ironclad. OpenAI is asking a small number of vendors to contribute hard examples, knowledgeable staff, secure test environments and research-usable data. If agents become more capable inside specialized products, vendors may gain a new way to make their services useful. They may also face a shift in how customers interact with them: people could issue instructions to an agent instead of using the product’s screens. The vendor’s underlying rules, records, controls and audit trail would then carry more weight than its interface.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Built

The October 6 post described a research setup rather than a general product launch. OpenAI and Ironclad chose tasks that reflect work performed in contract-management software, then assessed model outputs against task-specific rubrics. The test covered 11 selected workflows; it does not establish how often agents can complete the full range of work that Ironclad customers perform.

OpenAI described the model training as practice in hosted copies of Ironclad’s product, using synthetic tasks built from public SEC-filed contracts. The company said it did not use non-public Ironclad customer data. The distinction matters because the work involves sensitive business processes, and the post’s account of the data sources is a statement from OpenAI, not an independent audit described in the source material.

The post’s final section invites other software companies to work with OpenAI on tasks current agents cannot reliably complete. It asks prospective partners to bring concrete failure cases, subject-matter expertise, secure testing environments and data that can safely be used for research. The stated aim is to train and evaluate agents against real business rules and multi-step workflows, rather than only generic computer-use tasks.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Results

The post does not establish how Astra would perform across the full range of Ironclad workflows, on live customer accounts or under production conditions. The test involved 11 research tasks, and the average criteria score does not identify which particular requirements were missed on each task. The source material also does not provide a breakdown of error severity or independent verification of the results.

It remains unclear whether the estimated time figures correspond to any real customer productivity gains. OpenAI explicitly characterized them as simulations, and the comparison does not establish whether an agent plus human review would be faster or less costly than an experienced person completing and checking the work. The source also does not describe a customer deployment, release timeline or commercial arrangement between the companies.

OpenAI’s statements about training data and the limits of customer data use are attributable to the company. The source material does not mention an independent audit of the data or test environment. Nor does it specify which other software companies may participate in the proposed partner program.

Amazon

AI-powered procurement workflow tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What OpenAI and Vendors May Test

OpenAI has invited a small number of software companies to bring difficult workflows to future research partnerships. The next developments to watch are whether other vendors join, what tasks they choose, how their results are scored and whether OpenAI publishes more detail on failure types and verification methods.

For any move from research to customer use, buyers will need clearer information about which rules an agent can handle, what happens when it misses a requirement, and who checks its work. The October 6 post describes training and evaluation work; it does not announce that customers can hand over contract workflows to Astra without supervision.

Amazon

automated NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad announce?

OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The post also invited other software companies to discuss similar research partnerships.

Does the 55% result mean Astra completed 55% of tasks?

No. OpenAI reported that Astra met an average 55% of the criteria in the task rubrics. That is not the percentage of tasks completed, and it does not show that each workflow was safe or ready for use without review.

Did Astra cut customer task times by nearly half?

The post does not report measured customer savings. OpenAI’s figures of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol are simulated estimates based on assumed processing and generation speeds.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. It stated that it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can companies use Astra to manage contracts now?

The source describes research and testing, not a general customer deployment. OpenAI’s post says human oversight remains necessary, and it does not provide a release timeline for unsupervised contract or procurement work.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Xiaomi MiMo V2.6

Xiaomi has released MiMo v2.6, a new version of its wireless communication technology, prompting increased industry attention amid limited confirmed details.

Mojo 1.0

MojoTech announced the release of Mojo 1.0, a new AI language model aimed at enterprise applications, with early access available now.

15 AI Image Generators For Exploring Creative Ideas In 2027

A 2027 guide compares 15 AI image generators and related prompt, commerce, and learning resources, while noting what buyers should verify.

10 Best Mini PCs For Local AI Projects To Explore In 2026

A 2026 mini PC roundup compares memory, storage and expansion options for local AI, but the supplied source identifies only seven models.