
Core Purpose
The core purpose of the ERA Project is to develop the theoretical and empirical foundations needed to assess - and enhance - error-reasoning in AI scientists and human-AI teams, with a focus on the life sciences.
Error Theory
We will construct a comprehensive error framework for the life sciences to categorise experimental failures and the cognitive strategies used to resolve them. This framework will build on existing literature in Philosophy of Error and our own observations of error reasoning in scientific practice.
Benchmarking
Our team will use this framework to create benchmarks that evaluate the ability of AI agents and human-AI teams to navigate the complexities of scientific error.
Autonomous Reasoning
By fine-tuning frontier models on our specialised error data, we hope to eventually enable the next generation of agents to conduct robust research, either on their own or as part of human-AI teams. As part of our wider research we will also analyse how automation is changing scientific research.
Background
Recent advances in artificial intelligence have fuelled the dream of building “AI scientists” – systems that can autonomously perform research and thus accelerate the scientific discovery process. To realise this dream the novel AI agents need the ability to engage in the kind of reasoning that scientists mostly engage in, namely, error-reasoning. Experiments often fail and the average researcher spends the majority of their time trying to figure out what has gone wrong and why.
It is, however, uncertain to what extent existing AI agents can reason about error. Underlying this uncertainty is the fact that scientific error-reasoning has not been widely or deeply datafied: scientists work through errors in weekly laboratory meetings, on whiteboards, or in the hallways of a conference venue. These discussions rarely find their way into published materials and are thus underrepresented in the data on which AI models are trained. In addition, this kind of experimental troubleshooting might display features that everyday instances of troubleshooting don't possess.
To address this uncertainty, we need benchmarks that allow us to assess the extent to which AI models can reason about scientific error. However, even our most sophisticated benchmarks for AI agents don’t currently test for this type of reasoning. Instead, they focus on abilities such as the retrieval of textbook knowledge or on individual research tasks, for instance the extraction of information from a figure in a paper.
Developing effective benchmarks requires a good understanding of error-reasoning in science: what are the types of errors scientists encounter and what are the strategies they usually deploy to address them? Unfortunately, we still lack a systematic and comprehensive theory of error in science. This means that we are not yet in a good position to build the benchmarks we need.

Project Structure
Project start date: 1 September 2026
Duration: 4 years
Funder: Leverhulme Trust (Research Leadership Award)
The ERAs Project brings together an interdisciplinary team of researchers to address the problem of AI-driven automation and error-reasoning from the bottom up. In the first phase we will develop an error framework for a specific scientific discipline, the life sciences (ErrorTheory). The narrow focus on one discipline will allow us to build a systematic and comprehensive account of error for a field that is currently a focal point for the development and deployment of AI Scientists. In the second phase our team will use this ErrorTheory to build a systematic dataset which captures how researchers in the biology laboratory deal with different types of error (ErrorData). In the third phase we will use ErrorData to construct novel benchmarks that can assess the error-reasoning ability of existing and future AI agents, as well as human-AI teams (ErrorBench).
Looking into the future, ErrorData will also allow us to develop novel Error-Reasoning Agents (ERAs) through supervised finetuning of frontier models. Ultimately, our project will help build the foundations for the careful and effective introduction of AI agents into the research process.
Besides this focus on error-reasoning and benchmark development, our team will also study how automation, and in particular AI-driven automation, is changing the way research is performed.
