🔍 Read the full analysis: Three Things To Consider About Ironclad And OpenAI Agent Training on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training a frontier model in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task rubric criteria, while OpenAI’s estimated completion times were simulations rather than measured customer savings. The work also marks an invitation for other software companies to help train agents on difficult workflows.
OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on professional workflows inside hosted copies of contract-management software company Ironclad’s product. Across 11 legal, commercial and procurement tasks, Astra met an average 55% of rubric criteria, according to OpenAI; the company says its time figures are simulated estimates, not measured savings for customers.
The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s chosen jurisdiction. OpenAI estimates that an experienced user would take 30 to 40 minutes on each task.
Tasks were evaluated against 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. An internal model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria. These figures describe rubric performance, not the percentage of tasks completed successfully.
OpenAI’s estimated time per attempt was 19.2 minutes for Astra, versus 37 minutes for GPT-5.6 Sol. The post says these are simulated estimates based on assumed processing and generation speeds. They are not observations of customers using the system and do not establish real-world productivity gains. OpenAI said it used synthetic tasks built from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Agents Face Contract Workflow Controls
The findings show why performance in business software cannot be judged by a single average score or a faster estimated run time. In contract and procurement work, missing one rule can invalidate the outcome: a process may need Finance approval above a spending limit, Security review for particular requests and Legal review for nonstandard terms. A workflow that overlooks one required check may create a control failure, even if it satisfies many other criteria.
OpenAI’s reported 55% average criteria score is a notable research result, but it does not show that Astra is ready to handle contract processes without oversight. The study’s rubric-based results and simulated timing estimates leave practical questions about accuracy, error severity and review costs. For businesses, the immediate issue is not simply whether an agent can operate their software, but whether it reliably follows the rules that make the workflow safe.
The work also has implications for software companies. Training agents in a vendor’s product could make that product more useful, while giving the vendor evidence about where automated systems fail. But if customers increasingly give instructions to agents rather than use screens themselves, a software company’s durable value may depend more on its business rules, records, audit trails and controls than on its interface.
contract management software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Product Use to Model Training
Ironclad is a contract-management software company, not the name of a new agent framework. The OpenAI post describes a project in which models practised tasks inside hosted copies of Ironclad’s product. OpenAI framed the goal as training models to understand business rules, carry out multi-step work in specialised software and check that finished work meets the original requirements.
The project used tasks developed with people familiar with the work and synthetic training material derived from public SEC-filed contracts. OpenAI says it excluded non-public Ironclad customer data, OpenAI customer data and OpenAI internal contracts. The post also invites a small number of other software companies to partner on workflows current agents cannot reliably complete. It asks prospective partners to provide a concrete failure case, subject-matter experts, a secure test environment and data suitable for research.
This places the Ironclad work within a broader effort to train agents on real professional tasks rather than general computer use alone. The published account describes a research collaboration and results on a limited set of tasks; it does not announce a general release of autonomous contract-processing capabilities.
As an affiliate, we earn on qualifying purchases.
Deployment Readiness Remains Unproven
The published results do not establish how Astra would perform across a wider range of contracts, companies or unusual cases. OpenAI reported average criteria coverage and one showcase task score, but the post’s summary does not supply enough detail to determine which requirements were missed on each task or how serious those failures were. The 55% figure is not a task-completion rate, and a high score on one example does not show consistent performance across the set.
It is also unclear whether the reported results have been independently replicated or how performance changes when an agent is asked to work with live customer information and production systems. OpenAI says the training used public-contract-derived synthetic tasks and excluded specified customer and internal data, but the account does not establish what commercial deployment arrangements, if any, will follow. The simulated time estimates do not answer whether human review would erase possible time savings.
OpenAI’s own description acknowledges the need for human oversight when an agent may lose track of a business rule. Until error rates, missed-control patterns and review requirements are described in greater detail, the results should be read as research performance on a narrow task set, not evidence that businesses can safely delegate contract workflows end to end.
AI-powered contract review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Further Software Partnerships Ahead
OpenAI says it is inviting a small number of software companies to work on tasks agents do not yet complete reliably. The proposed partners would need to bring specific examples of failure, people with deep knowledge of the work, a secure test environment and data that can be used safely for research. The company has not, in the source account, named additional partners or published a schedule for further results.
For companies considering agent access to contract, finance or customer systems, the next practical step is to ask vendors how they measure workflow performance. Buyers should seek task-level results, details on failed criteria, safeguards for approvals and a clear account of when a person must review or approve an action. They should also ask what data is used in training and testing, and how actions are recorded for audit.
Further evidence will be needed to show whether models can meet all required controls reliably and whether they reduce total work once human checking is included. Until then, OpenAI and Ironclad’s results point to a training approach under development, with human review still part of the process.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad announce?
OpenAI described training models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The post also invites other software companies to collaborate on difficult agent workflows.
What does GPT-6 Astra’s 55% score mean?
It is the average share of rubric criteria met across the tasks, according to OpenAI. It is not the share of tasks completed successfully, and it does not mean the agent is ready to run workflows without review.
Did Astra save customers time?
The source does not report measured customer savings. OpenAI’s 19.2-minute estimate is simulated, based on assumed processing and generation speeds, rather than observed use by customers.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Can businesses use agents for contract work without human review?
The published results do not support that conclusion. OpenAI’s account recognizes that agents can lose track of business rules, and the reported average score leaves material questions about missed requirements. Human oversight remains necessary based on the information provided.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
