Reasoning with Machines Lab · Oxford

The science of AI systems that earn trust.

Led by Prof. Adam Mahdi, our lab advances the science of AI evaluation, benchmarking, safety and security. Through rigorous empirical research, we study how LLMs and agentic systems reason, interact with humans and drive scientific discovery.

Established atOxford Internet Institute
Research registerEvaluation · Safety · Agentic AI · Human-AI
Recent venuesNature Medicine · ICML · NeurIPS · ICLR
Featured research

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

A benchmark that tells real navigation apart from stochastic search when agents work over document collections.

AuthorsŁ Borchmann, J Van Landeghem, M Turski, S Padarha, RO Kearns, A Mahdi, et al.
ICML (Spotlight)May 2026Benchmarks and Evaluation
Read the paper
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Plate · Document collection navigationOpen access · arXiv
Research Themes

Four standing public commitments.

OxRML's research is organised around four long-running pillars. They guide what we publish, what we open-source, and what we will and won't take money to do. We update them publicly when the science demands it.

01Evaluation

Benchmarks and Evaluation

We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.

What this commits us to
02Safety

AI Safety and Security

From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.

What this commits us to
03Agentic

Agentic AI for Science

We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.

What this commits us to
04Human-AI

Human–AI Interaction

We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.

What this commits us to
Prof. Adam Mahdi
Plate · Principal Investigator2026
From the principal investigator
Adam leads OxRML. The group studies how language models reason, how people work with them, and how agentic systems behave on real scientific and decision-making tasks.
Prof. Adam MahdiLab Lead · Associate Professor · Oxford Internet Institute, University of Oxford
The lab

15 researchers, one programme.

DPhils · MScs · Visiting Fellows · AffiliatesSee all
Felix Krones
Felix KronesDPhil StudentMultimodal AI, digital health
Djavan De Clercq
Djavan De ClercqDPhil StudentAI and food security, LLMs
Andrew M. Bean
Andrew M. BeanDPhil StudentLLM evaluations, human–LLM interaction
Yushi Yang
Yushi YangDPhil StudentLLM & agentic post-training, AI alignment
Harry Mayne
Harry MayneDPhil StudentLLM interpretability, AI safety, LLM evaluations
Jessica Rodrigues
Jessica RodriguesDPhil StudentKnowledge graphs, metascience
Guy Parsons
Guy ParsonsDPhil StudentHealthcare AI, digital health
Karolina Korgul
Karolina KorgulDPhil StudentAI safety, agentic AI
Ryan Othniel Kearns
Ryan Othniel KearnsDPhil StudentScience of evals, reasoning in LLMs
Shreyansh Padarha
Shreyansh PadarhaDPhil StudentAI for science, AI safety, LLM evaluations
Mia Kussman
Mia KussmanMSc StudentHuman–LLM interaction, LLM evaluations
Caleb Tan
Caleb TanMSc StudentLLM evaluations, reasoning
Sebastian Petric
Sebastian PetricVisiting Policy FellowLLMs and financial time series
Tristan Naidoo
Tristan NaidooResearch AffiliatePublic health AI, LLM evaluations
Josh Lawman
Josh LawmanEntrepreneur in ResidenceResearch-to-product translation
Engage with the lab

Three ways to work with us.

We work with industry, foundations, and government on the questions our research touches. Every engagement keeps the science we publish honest.

hello@oxrml.com
Programme · 01

Workshops for industry teams

On-site sessions for product and ML teams on evaluation, safety, and agent reliability.

Half-day to multi-week formats. For teams shipping LLM products in healthcare, finance, retail, and government.

Book a workshop
Programme · 02

Tools co-built with engineering partners

We work with engineering partners to turn lab work into tools other teams can run.

Evaluation harnesses, safety dashboards, agentic-research platforms. We build them with partners we trust, carrying the research methods through to the code.

See our builds
Programme · 03

Research partnerships

Applied research collaborations with foundations, governments, and large companies.

Multi-year programmes: shared roadmaps, sponsored DPhil studentships, named labs.

Start a conversation