
Faraday says it found research taste. The benchmark still has a human-shaped hole.
Inherent's Faraday shows an intriguing way to train AI to choose experiments, but its public story still leaves access, pricing, research-data handling, and human sign-off outside the product.
"I got curious about this, and I went off and I did these experiments." 1
That is Edward Hughes describing his ideal AI teammate. It is also a useful description of Faraday: an agent that wanders through a research problem, chooses what to try, and returns with results instead of waiting for a prompt-shaped assignment.
On August 22, 2026, TechCrunch reported that London AI lab Inherent had released Faraday, a research agent built to reproduce findings from published scientific papers without being given the answer first. 1
The pitch sounds like an AI scientist. The public evidence describes a much narrower product: an automated research trainee with an interesting way to choose experiments, a frontier coding tool in its backpack, and a benchmark that leaves the buying question unanswered.
What Faraday actually does
Faraday starts with a published paper and tries to reproduce its findings independently. The agent has to work out which experiments are worth running, design those experiments, use code to run them, and compare the outcome with the paper's claims. Inherent calls the judgment behind those choices "research taste." 1
The training choice is reinforcement learning. Inherent rewards useful outcomes instead of spelling out every procedure in advance. The bet is that this will teach an agent how to make better experimental choices across scientific fields, rather than making it good at following one fixed recipe. 1
The model stack makes the product more interesting and the headline less tidy. Faraday runs on Qwen 3.6, a 27-billion-parameter model. Inherent uses OpenAI's GPT-5.5 Codex for coding instead of building its own coding tool. 1
That makes Faraday a workflow rather than a single small model pulling a rabbit out of a beaker. The research policy comes from Inherent's agent, while a separate frontier coding tool handles an important part of execution. The useful comparison is therefore between complete research workflows, not model size alone.

The benchmark is doing most of the selling
Inherent says Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 on independently reproducing published research. The comparison covers a specific task, and the company says it cared about more than accuracy: Faraday also had to show judgment about which experiments to run and how to design them. 1
That is a good research question. It is also a very convenient product demo. The report gives the rival model names and the winning claim, while leaving out the score table, paper count, task selection, compute budget, harness configuration, and independent validation that would let a buyer judge the result. 1
The missing detail changes the meaning of "outperformed." A benchmark win can show that an agent learned a useful behavior under a defined test. It cannot tell a lab how often the agent will choose a meaningful experiment, handle private research material, or produce a result that a human can sign off.
A research team still needs answers to four practical questions:
- Can a team give Faraday its own papers and receive a reproducible result?
- Which parts of the workflow require a human to translate a paper into runnable code?
- How often does the agent choose an experiment that tests the claim rather than merely producing a plausible-looking chart?
- What does a failed run cost in compute, review time, and researcher attention?
The TechCrunch report gives readers no signup route, API terms, plan, subscription, or price. It calls Faraday released while leaving the commercial door closed. 1
The access layer is still a blank page
The implied audience is a scientific researcher or research team. Faraday's starting point is a published paper, and Hughes compares the task with the way PhD students learn research by reproducing earlier work. 1
For that audience, the product boundary matters more than the phrase "AI scientist." Publicly described facts currently look like this:
- Mechanics: Faraday uses reinforcement learning to develop better choices about experiments while reproducing paper findings. 1
- Model stack: Faraday runs on Qwen 3.6 with 27 billion parameters and delegates coding to GPT-5.5 Codex. 1
- Access: Inherent describes Faraday as released. The report supplies no public access instructions or deployment route. 1
- Pricing: The report gives no price, subscription tier, API rate, or usage meter. 1
- Data and permissions: Published papers are the described research input. The report gives no upload workflow, paper-licensing guidance, retention policy, or customer-data terms. 1
For a lab, the last line carries the largest practical risk. A replication may touch code, datasets, software versions, private credentials, licensed material, and domain judgment. Faraday's public description explains the research ambition and training method. The handling of that research remains undisclosed.
The old idea got a new optimizer
Paper replication is already part of scientific training. Hughes told TechCrunch that many PhD students begin by reproducing earlier work. Faraday's fresh idea is the attempt to make the experiment-selection habit trainable through reinforcement learning, then place that habit inside an agent that can use a coding tool. 1
The design has a clear causal chain:
The problem: An AI that contributes to science must do more than turn a question into fluent prose. Inherent wants an agent that can help discover new knowledge across scientific fields. 1
The constraint: Researchers cannot write a complete rulebook for every useful experiment. The valuable choice often arrives before the code: which variable to change, which control to add, and which result is worth checking next.
The design choice: Inherent uses reinforcement learning to reward outcomes, gives Faraday a relatively small base model, and supplies GPT-5.5 Codex as the coding specialist. 1
The failure mode: The reward target can become the product. If the benchmark rewards a convincing reproduction, an agent may learn to optimize the visible score while the human researcher still has to decide whether the experiment was scientifically meaningful. The report gives readers too little information to separate those outcomes. 1
The consequence: Faraday has demonstrated an intriguing research workflow claim. It has yet to demonstrate a research service that a lab can price, govern, connect to private material, and trust with an important result.
Verdict
Faraday is a promising research demo wrapped in a product-shaped headline. The useful idea is specific: train an agent to choose experiments and use an existing coding specialist instead of asking a single model to perform the whole scientist costume. The purchasing case is far less mature. Inherent has disclosed the ambition, the reinforcement-learning bet, and a claimed win on paper replication, while leaving access, pricing, research-data handling, benchmark detail, and human sign-off outside the public story. Research teams should treat Faraday as a signal that experiment selection is becoming an AI product target, not as a ready-made lab colleague. Until Inherent publishes the door, the meter, and the audit trail, the most autonomous thing about Faraday is the marketing copy.
References
- 1
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Plaud One puts an agent in your earbuds. The cellular case is the product.
- Perplexity Portable Computer puts the agent on a $4,699 desk. The box is the product.
- Keenable indexed the web for agents. The turnstile is the product.
- Thomson Reuters built its legal model. The product is still CoCounsel.
- ChatGPT for Teens puts a safety gate around the chatbot. The gate is the product.
- Warp Factories promises a software factory. The approval queue is still human.
- Ramp Router promises cheaper model bills. The toll booth keeps the data.
- Google gave students a free AI study buddy. The syllabus is the onboarding form.
