OxRML · University of OxfordOxford Internet Institute

Reasoning with Machines Lab

@ University of Oxford

Led by Prof. Adam Mahdi, we work on the science of evaluating, benchmarking and securing modern AI. Our empirical research asks how LLMs and agentic systems reason, collaborate with humans and accelerate scientific discovery — alongside industry partners deploying these systems where reliability matters.

Research themes04
01
Benchmarks & Evaluation
02
AI Safety & Security
03
Agentic AI for Science
04
Human–AI Interaction
Benchmarks and EvaluationAI Safety and SecurityAgentic AI for ScienceHuman–AI InteractionMechanistic interpretabilityRed-teamingCapability elicitationBenchmark designBias & toxicityDomain-grounded agentsScientific discoveryAI governanceDistributional robustnessSelf-explanationReasoning under pressureBenchmarks and EvaluationAI Safety and SecurityAgentic AI for ScienceHuman–AI InteractionMechanistic interpretabilityRed-teamingCapability elicitationBenchmark designBias & toxicityDomain-grounded agentsScientific discoveryAI governanceDistributional robustnessSelf-explanationReasoning under pressure
01·What we do

Four research lines, asking whether we can trust what these systems do next.

01 / 04

Benchmarks and Evaluation

We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.

  • Benchmark design
  • Statistical evaluation
  • Capability elicitation
  • Contamination audits
02 / 04

AI Safety and Security

From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.

  • Mechanistic interpretability
  • Red-teaming
  • Agentic misalignment
  • Policy translation
03 / 04

Agentic AI for Science

We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.

  • Literature synthesis
  • Hypothesis generation
  • Evidence grounding
  • Domain transfer
04 / 04

Human–AI Interaction

We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.

  • Field experiments
  • Decision-aid design
  • Clinical evaluation
  • Policy translation
02·Selected work

Recent publications

All publications
03·For industry

We help teams ship AI they can defend.

Two ways to work with us: third-party evaluation of your models and agents, or a focused engagement that turns one of our research outputs into a tool you own.

Evaluation
Pre-deployment audits, custom benchmarks, agentic red-teaming.
Co-build
We work with engineering partners to turn lab work into tools other teams can run.
Sectors
SaaS, public sector, financial services, healthcare.
Engagement
12–24 weeks · NDA-friendly · publishable outcomes negotiable.
04·The lab

DPhils, MScs, and visiting fellows.

Full team
Felix KronesFK
Felix Krones
DPhil Student
Djavan De ClercqDDC
Djavan De Clercq
DPhil Student
Andrew M. BeanAMB
Andrew M. Bean
DPhil Student
Yushi YangYY
Yushi Yang
DPhil Student
Harry MayneHM
Harry Mayne
DPhil Student
07·Join the lab

We welcome applications from motivated students.

We accept students through Oxford Internet Institute graduate programmes. Applications are typically due in early January for September entry.

05·Lab notes

Recent activity

All notes
May 2026
Paper
Papers accepted at ICML 2026!
April 2026
Conference
OxRML at ICLR 2026
February 2026
Award
Ryan Othniel Kearns Wins MSc Thesis Prize
February 2026
Paper
New Paper in Nature Medicine!
December 2025
Conference
OxRML @ NeurIPS 2025
© 2026 OxRML · University of Oxford
v.2026.05