Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication
The company’s result centers on reproducing known findings, not making discoveries. Its next test is whether that training approach can move from replication to reliable new science.
Listen to this story
The audio brief
Story brief
3 key pointsInherent’s Faraday is a research-agent system built around a 27B Qwen 3.6 model and reinforcement learning for experimental choices. The London startup claims it outperformed Claude Opus 4.8 and GPT-5.5 on unpublished details of a replication evaluation, but provided no scores, papers, or methodology. Because Faraday delegates coding to OpenAI’s GPT-5.5 Codex, the result measures system design rather than a...
- 01
Faraday reproduces existing paper findings; it has not yet demonstrated new scientific discovery.
- 02
Inherent disclosed no scores, test papers, or evaluation methodology, limiting independent assessment of the claimed win.
- 03
The system combines Qwen 3.6, reinforcement learning, and GPT-5.5 Codex for coding.
Inherent says its Faraday agent independently reproduced findings from published scientific papers without receiving the answers in advance, outperforming Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 in the company’s evaluation. That comparison puts a young London lab’s bet on research-agent design against much larger frontier systems.
A rehearsal for the larger goal
Faraday was asked to recover findings already reported in papers, rather than generate new scientific knowledge. That distinction is central: Inherent’s stated long-term goal is an agent that contributes to discovery across scientific fields. Cofounder and chief scientist Edward Hughes described replication as a common starting exercise for researchers, saying many PhD students begin there.
The company’s approach is built around what it calls research taste: deciding which experiments are worth running and how to design them. Inherent used reinforcement learning, which rewards desired outcomes, to train Faraday toward that experimental judgment. It is betting this method will generalize across fields more effectively than training an agent primarily on descriptions of scientific practice.
A smaller core model, with outside tools
Inherent chose not to build its own coding tool, instead having Faraday use GPT-5.5 Codex. That design choice makes the reported win less a clean base-model contest than a test of how a research agent combines a core model, reward-based training and specialized software.
The disclosed Faraday setup
- Qwen 3.6 as the 27-billion-parameter core model described by Inherent.
- Reinforcement learning aimed at experimental selection and design.
- GPT-5.5 Codex for coding instead of an Inherent-built coding tool.
The evidence still stops at replication
The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison. It supports Inherent’s early case for its research-agent approach, but does not yet demonstrate that Faraday can produce reliable new science.
Inherent emerged from stealth with a $50 million seed round and has about a dozen employees. It plans to grow to roughly 20 to 25 people by year-end; its most consequential next proof point will be whether experimental judgment carries beyond reproducing published results.
Sources
- techcrunch.comInherent, founded by DeepMind alumni, says its AI 'teammate' just outperformed Anthropic and OpenAI at replicating research | TechCrunch