Inherent, an AI lab based in London, has released Faraday, an agent built specifically for reproducing published research. It sits on top of a 27-billion-parameter model, yet on this task the published results put it ahead of the far larger Claude Opus 4.8 and GPT-5.5. The point was never to top a leaderboard, but to teach an agent how a working scientist actually behaves.
Redrawing a figure without seeing the answer
Faraday was trained on Replica, a space of reinforcement learning tasks. The agent receives a research paper and is asked to reproduce one of its figures. It never sees the original plot, and both its time and its compute are capped. All it has to work from is the text and the method described in it.
The initial suite covers 310 tasks drawn from 100 papers in machine learning and AI for science. The domains range from natural language processing to materials science and weather forecasting, so tuning for the quirks of a single field does not get an agent very far.
Reproducing a figure sounds like busywork, but Inherent argues that the core of research ability lives right there. A paper records the steps that worked. Everything the authors tried and discarded on the way never makes it onto the page. To pull off a replication, an agent has to reconstruct that missing trial and error on its own, which demands the hypothesis-driven exploration that defines research in the first place.
A 27B model ahead of far larger rivals
The comparisons were Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5, run in the Claude Code and Codex harnesses respectively with thinking effort set to the highest level. Faraday's base is the 27-billion-parameter version of Qwen 3.6, which is not in the same weight class at all.
Inherent reports that Faraday produced more faithful replications in every category of the suite, with the widest margins in meta-learning, structural biology and materials science. It also degrades less on recent research, which the company reads as evidence that the agent applies learned procedures to material its base model never saw during pretraining.
Grading research taste
Scoring was the hard part. A plot that looks the same is not the same thing as a successful replication. Experimental design, sound scientific practice, faithfulness to the original claims and sensible use of limited resources all have to count. Inherent calls this bundle research taste.
The lab put an LLM in the judge's seat and ran a human study to check how closely its verdicts track expert judgement. The catch is that a language model judge is stochastic, and that variance turns straight into reward noise. The fix was a per-task rubric, which proved more consistent than a generic judge and agreed more often with human raters. During training the team layered on two more measures: aggregating several judgements per sample, and assigning credit turn by turn rather than only at the end.
A small model directing a much bigger one
The other striking detail is that Faraday has no coding tool of its own. The actual code is written by GPT-5.5 Codex, while Faraday decides what to test and how. It is the same division of labour a human researcher has with existing software, transplanted into an agent.
More surprising still, Faraday was trained against GPT-5.4-mini yet performed better when handed the more capable GPT-5.5 Codex at test time. It steers a model orders of magnitude larger than itself and improves the outcome. If coding agents keep getting stronger, Inherent's bet is that the judgement directing them becomes more valuable, not less. Faraday also runs without a hand-coded evolutionary harness and without any test-time reward, which sets it apart from earlier automated research agents.
Inherent is young. It came out of stealth in May 2026 with a 50 million USD seed round (about 8 billion yen), founded by Google DeepMind alumni and based in the King's Cross district of London. The team is around 12 people today, with plans to reach 20 to 25 by the end of the year. Faraday is its first public deliverable.
※1 USD = 159 JPY (as of August 21, 2026)
Summary
Faraday runs on a 27-billion-parameter base and still outperformed Claude Opus 4.8 and GPT-5.5 at replicating research papers. Three choices carried it: a task design that hides the answer and asks for the figure back, a rubric-based judge that grades the quality of the science rather than pixel similarity, and a division of labour that leaves implementation to a much larger model. Read as an argument, it says that scaling the model is not the only way to make AI useful to science.
