Modelspublished

Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication

The company’s result centers on reproducing known findings, not making discoveries. Its next test is whether that training approach can move from replication to reliable new science.

By 2 min read
Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication

Listen to this story

The audio brief

About 1:24
0:001:24
Read transcript
Inherent says its Faraday research agent has beaten Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at reproducing findings from published papers. But the result is narrower than that headline sounds: Faraday was asked to recover known results, not make a new discovery. The London startup says Faraday did the work without seeing answers in advance. Its core is Qwen 3.6, a model with 27 billion parameters, trained with reinforcement learning to make better choices about which experiments to run and how to design them. Inherent calls that capability “research taste”—the practical judgment researchers use to decide what is worth testing. There is an important wrinkle in the comparison. Faraday uses GPT-5.5 Codex for coding, rather than a coding tool built by Inherent. So this is a test of a complete research-agent system: a smaller core model, reward-based training, and outside software—not a clean contest between Qwen and frontier models. And the evidence is thin. Inherent has not released scores, the test papers, or its evaluation method, so the claimed win cannot be independently assessed. The company has raised 50 million dollars in seed funding and has about 12 employees. The decisive question is whether this experimental judgment can carry from replication into reliable new science.

Story brief

3 key points

Inherent’s Faraday is a research-agent system built around a 27B Qwen 3.6 model and reinforcement learning for experimental choices. The London startup claims it outperformed Claude Opus 4.8 and GPT-5.5 on unpublished details of a replication evaluation, but provided no scores, papers, or methodology. Because Faraday delegates coding to OpenAI’s GPT-5.5 Codex, the result measures system design rather than a...

  1. 01

    Faraday reproduces existing paper findings; it has not yet demonstrated new scientific discovery.

  2. 02

    Inherent disclosed no scores, test papers, or evaluation methodology, limiting independent assessment of the claimed win.

  3. 03

    The system combines Qwen 3.6, reinforcement learning, and GPT-5.5 Codex for coding.

Inherent says its Faraday agent independently reproduced findings from published scientific papers without receiving the answers in advance, outperforming Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 in the company’s evaluation. That comparison puts a young London lab’s bet on research-agent design against much larger frontier systems.

A rehearsal for the larger goal

Faraday was asked to recover findings already reported in papers, rather than generate new scientific knowledge. That distinction is central: Inherent’s stated long-term goal is an agent that contributes to discovery across scientific fields. Cofounder and chief scientist Edward Hughes described replication as a common starting exercise for researchers, saying many PhD students begin there.

The company’s approach is built around what it calls research taste: deciding which experiments are worth running and how to design them. Inherent used reinforcement learning, which rewards desired outcomes, to train Faraday toward that experimental judgment. It is betting this method will generalize across fields more effectively than training an agent primarily on descriptions of scientific practice.

A smaller core model, with outside tools

Inherent chose not to build its own coding tool, instead having Faraday use GPT-5.5 Codex. That design choice makes the reported win less a clean base-model contest than a test of how a research agent combines a core model, reward-based training and specialized software.

The disclosed Faraday setup

  • Qwen 3.6 as the 27-billion-parameter core model described by Inherent.
  • Reinforcement learning aimed at experimental selection and design.
  • GPT-5.5 Codex for coding instead of an Inherent-built coding tool.

The evidence still stops at replication

The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison. It supports Inherent’s early case for its research-agent approach, but does not yet demonstrate that Faraday can produce reliable new science.

Inherent emerged from stealth with a $50 million seed round and has about a dozen employees. It plans to grow to roughly 20 to 25 people by year-end; its most consequential next proof point will be whether experimental judgment carries beyond reproducing published results.

Sources

  1. techcrunch.comInherent, founded by DeepMind alumni, says its AI 'teammate' just outperformed Anthropic and OpenAI at replicating research | TechCrunch