Faraday says it found research taste. The benchmark still has a human-shaped hole.

Faraday says it found research taste. The benchmark still has a human-shaped hole.

Inherent's Faraday shows an intriguing way to train AI to choose experiments, but its public story still leaves access, pricing, research-data handling, and human sign-off outside the product.

"I got curious about this, and I went off and I did these experiments." 1
That is Edward Hughes describing his ideal AI teammate. It is also a useful description of Faraday: an agent that wanders through a research problem, chooses what to try, and returns with results instead of waiting for a prompt-shaped assignment.
On August 22, 2026, TechCrunch reported that London AI lab Inherent had released Faraday, a research agent built to reproduce findings from published scientific papers without being given the answer first. 1
The pitch sounds like an AI scientist. The public evidence describes a much narrower product: an automated research trainee with an interesting way to choose experiments, a frontier coding tool in its backpack, and a benchmark that leaves the buying question unanswered.

What Faraday actually does

Faraday starts with a published paper and tries to reproduce its findings independently. The agent has to work out which experiments are worth running, design those experiments, use code to run them, and compare the outcome with the paper's claims. Inherent calls the judgment behind those choices "research taste." 1
A simplified four-stage view of Faraday's paper-replication task
Self-made explanatory diagram based on Inherent's description of Faraday's paper-replication task. It is not a Faraday screenshot. 1
The training choice is reinforcement learning. Inherent rewards useful outcomes instead of spelling out every procedure in advance. The bet is that this will teach an agent how to make better experimental choices across scientific fields, rather than making it good at following one fixed recipe. 1
The model stack makes the product more interesting and the headline less tidy. Faraday runs on Qwen 3.6, a 27-billion-parameter model. Inherent uses OpenAI's GPT-5.5 Codex for coding instead of building its own coding tool. 1
That makes Faraday a workflow rather than a single small model pulling a rabbit out of a beaker. The research policy comes from Inherent's agent, while a separate frontier coding tool handles an important part of execution. The useful comparison is therefore between complete research workflows, not model size alone.
Four Inherent cofounders standing and sitting in front of a whiteboard covered with research notes and formulas
From left to right, TechCrunch identifies the Inherent cofounders as Louis Kirsch, Kaloyan Aleksiev, Tantum Collins, and Edward Hughes. The whiteboard is a more honest product diagram than a glowing chatbot window: the work depends on experiments, tools, and judgment. 1

The benchmark is doing most of the selling

Inherent says Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 on independently reproducing published research. The comparison covers a specific task, and the company says it cared about more than accuracy: Faraday also had to show judgment about which experiments to run and how to design them. 1
That is a good research question. It is also a very convenient product demo. The report gives the rival model names and the winning claim, while leaving out the score table, paper count, task selection, compute budget, harness configuration, and independent validation that would let a buyer judge the result. 1
The missing detail changes the meaning of "outperformed." A benchmark win can show that an agent learned a useful behavior under a defined test. It cannot tell a lab how often the agent will choose a meaningful experiment, handle private research material, or produce a result that a human can sign off.
A research team still needs answers to four practical questions:
  • Can a team give Faraday its own papers and receive a reproducible result?
  • Which parts of the workflow require a human to translate a paper into runnable code?
  • How often does the agent choose an experiment that tests the claim rather than merely producing a plausible-looking chart?
  • What does a failed run cost in compute, review time, and researcher attention?
The TechCrunch report gives readers no signup route, API terms, plan, subscription, or price. It calls Faraday released while leaving the commercial door closed. 1

The access layer is still a blank page

The implied audience is a scientific researcher or research team. Faraday's starting point is a published paper, and Hughes compares the task with the way PhD students learn research by reproducing earlier work. 1
For that audience, the product boundary matters more than the phrase "AI scientist." Publicly described facts currently look like this:
  • Mechanics: Faraday uses reinforcement learning to develop better choices about experiments while reproducing paper findings. 1
  • Model stack: Faraday runs on Qwen 3.6 with 27 billion parameters and delegates coding to GPT-5.5 Codex. 1
  • Access: Inherent describes Faraday as released. The report supplies no public access instructions or deployment route. 1
  • Pricing: The report gives no price, subscription tier, API rate, or usage meter. 1
  • Data and permissions: Published papers are the described research input. The report gives no upload workflow, paper-licensing guidance, retention policy, or customer-data terms. 1
For a lab, the last line carries the largest practical risk. A replication may touch code, datasets, software versions, private credentials, licensed material, and domain judgment. Faraday's public description explains the research ambition and training method. The handling of that research remains undisclosed.

The old idea got a new optimizer

Paper replication is already part of scientific training. Hughes told TechCrunch that many PhD students begin by reproducing earlier work. Faraday's fresh idea is the attempt to make the experiment-selection habit trainable through reinforcement learning, then place that habit inside an agent that can use a coding tool. 1
The design has a clear causal chain:
The problem: An AI that contributes to science must do more than turn a question into fluent prose. Inherent wants an agent that can help discover new knowledge across scientific fields. 1
The constraint: Researchers cannot write a complete rulebook for every useful experiment. The valuable choice often arrives before the code: which variable to change, which control to add, and which result is worth checking next.
The design choice: Inherent uses reinforcement learning to reward outcomes, gives Faraday a relatively small base model, and supplies GPT-5.5 Codex as the coding specialist. 1
The failure mode: The reward target can become the product. If the benchmark rewards a convincing reproduction, an agent may learn to optimize the visible score while the human researcher still has to decide whether the experiment was scientifically meaningful. The report gives readers too little information to separate those outcomes. 1
The consequence: Faraday has demonstrated an intriguing research workflow claim. It has yet to demonstrate a research service that a lab can price, govern, connect to private material, and trust with an important result.

Verdict

Faraday is a promising research demo wrapped in a product-shaped headline. The useful idea is specific: train an agent to choose experiments and use an existing coding specialist instead of asking a single model to perform the whole scientist costume. The purchasing case is far less mature. Inherent has disclosed the ambition, the reinforcement-learning bet, and a claimed win on paper replication, while leaving access, pricing, research-data handling, benchmark detail, and human sign-off outside the public story. Research teams should treat Faraday as a signal that experiment selection is becoming an AI product target, not as a ready-made lab colleague. Until Inherent publishes the door, the meter, and the audit trail, the most autonomous thing about Faraday is the marketing copy.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel