Research · 22 Aug 2026 · 2 min read

Inherent's Faraday Agent Beats Larger Models at Research

Inherent's specialised AI agent beat OpenAI and Anthropic at replicating research, showing the power of smaller, focused agentic systems for builders.

What happened

Inherent, a London-based AI lab from DeepMind alumni, announced its new AI agent, Faraday. According to a TechCrunch report, Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at a specific task — independently reproducing the findings of published scientific papers.

The agent achieved this on a comparatively small model, Qwen 3.6, which has just 27 billion parameters. Inherent framed the task as a standard training exercise for human scientists, designed to test an agent's ability to reason and experiment.

How the room's reading it

The focus for AI practitioners isn't just the benchmark result. It's the architecture. Inherent's team is framing this as a win for specialised, agentic systems over monolithic models. They leaned on reinforcement learning to teach 'research taste' — an instinct for good experimental design — rather than just optimising for accuracy.

Developers are noting that Faraday uses other models like OpenAI's GPT-5.5 Codex as tools, much like a human scientist uses existing software. The consensus is that this points toward a future of smaller, specialised agents designed for complex intellectual work, not just another race for parameter count.

Sailfish's take

We think this is the right way to build. Chasing parameter counts is a game for giants. The real opportunity is in building sharp, specialised agents that use foundation models as a commodity resource — just as Inherent used GPT-5.5 for coding.

The claim about teaching 'research taste' with reinforcement learning is bold, and we're watching closely to see if it holds up beyond paper replication. For now, this validates a core belief we have: the next wave of value won't come from bigger models, but from smarter agents that can automate specific, high-value intellectual work. This is the stack to build on.

Our take — your read?

Be the first to weigh in.

Sources
— END OF DISPATCH — Research