OpenAI Publishes Ironclad Findings, With Astra Scoring 55% on Contracting Tasks
The company’s 11-task evaluation measures detailed requirements, not completed contracts. Its faster timing figures are simulated estimates rather than customer savings.
Loading page…
The company’s 11-task evaluation measures detailed requirements, not completed contracts. Its faster timing figures are simulated estimates rather than customer savings.
Listen to this story
OpenAI’s October 6 findings offer a benchmark for agents handling complex professional workflows, not evidence that such work can be delegated end to end: on 11 Ironclad tasks, GPT-6 Astra met an average 55% of task criteria, while an internal model used during its development reached 63.7%. The collaboration shows how software partners can turn approval logic into practice environments for reinforcement learning. But the limited task set and need to preserve business rules make human oversight a near-term constraint on deployment.
GPT-5.6 Sol averaged 41.6% on the same 11 tasks; the scores measure criteria met, not the share of jobs completed successfully.
Tasks covered workflows such as NDA setup, procurement approvals, and updating legal clauses for a selected jurisdiction.
Each task had 8 to 50 assessment criteria, and OpenAI estimates an experienced user would take 30 to 40 minutes per task.
OpenAI published findings on October 6, 2026, from a research collaboration with contracting-software company Ironclad, showing progress without claiming reliable automation. GPT-6 Astra averaged 55.0% on an evaluation of 11 complex workflows, versus GPT-5.6 Sol’s 41.6%, according to OpenAI. The scores measure how well agents met task requirements—not the percentage of jobs completed successfully.
Ironclad is OpenAI’s first software-company partner in an effort to train and evaluate agents on complex professional work. Ironclad employees and OpenAI staff who use its software helped select tasks across legal, commercial and procurement work. OpenAI estimates an experienced user would need 30 to 40 minutes per task, on average.
The tasks included:
A software-purchasing process illustrates the challenge. An agent must translate business requirements into an intake form, document templates and approval rules. If Finance must approve purchases above a spending threshold, the agent must configure that rule and check requests on both sides of the threshold. The test is whether the whole process follows the rules, not simply whether the agent completes individual steps.
Each task had 8 to 50 assessment criteria, depending on complexity. Ironclad supplied hosted software environments where models could practice. OpenAI built synthetic tasks around representative workflows and used reinforcement learning—practice guided by feedback—to improve performance. Astra is its first frontier model trained on Ironclad tasks.
OpenAI used each model’s highest-scoring reasoning setting for the comparison. The findings cover only the 11 research tasks, not every Ironclad workflow.
An internal model used during Astra’s development scored 63.7% on the tasks. OpenAI says it aims to bring those further gains to future models, without announcing a new release here.
OpenAI’s training-data disclosure says the simulated tasks came from public contracts in the SEC’s EDGAR database, filtered to remove personal information. It says neither training nor evaluation used OpenAI customer data, its internal contracts, or nonpublic Ironclad customer data or contracts.
OpenAI also stresses that human oversight remains necessary: losing a business rule midway through a task limits what a software company can confidently delegate. Ironclad’s potential benefit is a route to more capable underlying models for its products, while preserving the controls legal and business teams depend on.
Loading discussion...
Join the conversation
Explain which mistakes would outweigh an otherwise strong result.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.