AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Three Things To Consider About Ironclad And OpenAI Agent Training on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task rubric criteria, while OpenAI’s estimated completion times were simulations rather than measured customer savings. The work also marks an invitation for other software companies to help train agents on difficult workflows.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on professional workflows inside hosted copies of contract-management software company Ironclad’s product. Across 11 legal, commercial and procurement tasks, Astra met an average 55% of rubric criteria, according to OpenAI; the company says its time figures are simulated estimates, not measured savings for customers.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s chosen jurisdiction. OpenAI estimates that an experienced user would take 30 to 40 minutes on each task.

Tasks were evaluated against 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. An internal model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria. These figures describe rubric performance, not the percentage of tasks completed successfully.

OpenAI’s estimated time per attempt was 19.2 minutes for Astra, versus 37 minutes for GPT-5.6 Sol. The post says these are simulated estimates based on assumed processing and generation speeds. They are not observations of customers using the system and do not establish real-world productivity gains. OpenAI said it used synthetic tasks built from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.

At a glance
reportWhen: Published October 6; current results ar…
The developmentOpenAI published details of an Ironclad partnership in which models practised contract and procurement workflows inside hosted copies of the software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Agents Face Contract Workflow Controls

The findings show why performance in business software cannot be judged by a single average score or a faster estimated run time. In contract and procurement work, missing one rule can invalidate the outcome: a process may need Finance approval above a spending limit, Security review for particular requests and Legal review for nonstandard terms. A workflow that overlooks one required check may create a control failure, even if it satisfies many other criteria.

OpenAI’s reported 55% average criteria score is a notable research result, but it does not show that Astra is ready to handle contract processes without oversight. The study’s rubric-based results and simulated timing estimates leave practical questions about accuracy, error severity and review costs. For businesses, the immediate issue is not simply whether an agent can operate their software, but whether it reliably follows the rules that make the workflow safe.

The work also has implications for software companies. Training agents in a vendor’s product could make that product more useful, while giving the vendor evidence about where automated systems fail. But if customers increasingly give instructions to agents rather than use screens themselves, a software company’s durable value may depend more on its business rules, records, audit trails and controls than on its interface.

Amazon

contract management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Product Use to Model Training

Ironclad is a contract-management software company, not the name of a new agent framework. The OpenAI post describes a project in which models practised tasks inside hosted copies of Ironclad’s product. OpenAI framed the goal as training models to understand business rules, carry out multi-step work in specialised software and check that finished work meets the original requirements.

The project used tasks developed with people familiar with the work and synthetic training material derived from public SEC-filed contracts. OpenAI says it excluded non-public Ironclad customer data, OpenAI customer data and OpenAI internal contracts. The post also invites a small number of other software companies to partner on workflows current agents cannot reliably complete. It asks prospective partners to provide a concrete failure case, subject-matter experts, a secure test environment and data suitable for research.

This places the Ironclad work within a broader effort to train agents on real professional tasks rather than general computer use alone. The published account describes a research collaboration and results on a limited set of tasks; it does not announce a general release of autonomous contract-processing capabilities.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Deployment Readiness Remains Unproven

The published results do not establish how Astra would perform across a wider range of contracts, companies or unusual cases. OpenAI reported average criteria coverage and one showcase task score, but the post’s summary does not supply enough detail to determine which requirements were missed on each task or how serious those failures were. The 55% figure is not a task-completion rate, and a high score on one example does not show consistent performance across the set.

It is also unclear whether the reported results have been independently replicated or how performance changes when an agent is asked to work with live customer information and production systems. OpenAI says the training used public-contract-derived synthetic tasks and excluded specified customer and internal data, but the account does not establish what commercial deployment arrangements, if any, will follow. The simulated time estimates do not answer whether human review would erase possible time savings.

OpenAI’s own description acknowledges the need for human oversight when an agent may lose track of a business rule. Until error rates, missed-control patterns and review requirements are described in greater detail, the results should be read as research performance on a narrow task set, not evidence that businesses can safely delegate contract workflows end to end.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Software Partnerships Ahead

OpenAI says it is inviting a small number of software companies to work on tasks agents do not yet complete reliably. The proposed partners would need to bring specific examples of failure, people with deep knowledge of the work, a secure test environment and data that can be used safely for research. The company has not, in the source account, named additional partners or published a schedule for further results.

For companies considering agent access to contract, finance or customer systems, the next practical step is to ask vendors how they measure workflow performance. Buyers should seek task-level results, details on failed criteria, safeguards for approvals and a clear account of when a person must review or approve an action. They should also ask what data is used in training and testing, and how actions are recorded for audit.

Further evidence will be needed to show whether models can meet all required controls reliably and whether they reduce total work once human checking is included. Until then, OpenAI and Ironclad’s results point to a training approach under development, with human review still part of the process.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad announce?

OpenAI described training models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The post also invites other software companies to collaborate on difficult agent workflows.

What does GPT-6 Astra’s 55% score mean?

It is the average share of rubric criteria met across the tasks, according to OpenAI. It is not the share of tasks completed successfully, and it does not mean the agent is ready to run workflows without review.

Did Astra save customers time?

The source does not report measured customer savings. OpenAI’s 19.2-minute estimate is simulated, based on assumed processing and generation speeds, rather than observed use by customers.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can businesses use agents for contract work without human review?

The published results do not support that conclusion. OpenAI’s account recognizes that agents can lose track of business rules, and the reported average score leaves material questions about missed requirements. Human oversight remains necessary based on the information provided.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Small Business Guide To Putting AI To Work

An OpenAI page is titled “Helping small businesses put AI to work,” but available information does not establish what it describes or when it was published.

Make Your Buyer Skills Part Of The Business Search

A proposed marketplace workflow would match business-for-sale listings to buyers’ operating skills and test whether it improves qualified inquiries.

2026’S Leading AI Technologies: 15 Solutions For Smart Buyers

Explore the 15 leading AI solutions of 2026, highlighting innovative tools and their impact on various industries for informed purchasing decisions.

Onimusha: Way Of The Sword Climbing The Steam Charts

Onimusha: Way of the Sword has climbed to the top 10 on Steam, reaching a peak of over 43,000 players. The cause of this spike remains unconfirmed.