Reasoning with Machines Lab
@ University of Oxford

Led by Prof. Adam Mahdi, the Reasoning with Machines Lab advances the science of AI evaluation, benchmarking, safety and security. Through rigorous empirical research, we study how LLMs and agentic systems reason, interact with humans, and drive scientific discovery. We work with industry partners deploying AI where reliability matters.

Prof. Adam Mahdi
Prof. Adam Mahdi
Lab Lead · Associate Professor · Oxford Internet Institute, University of Oxford

Adam leads OxRML. The group studies how language models reason, how people work with them, and how agentic systems behave on real scientific and decision-making tasks.

Papers
10
open access
Themes
4
open access
Team
15
+ lead

Engage with OxRML

For individuals

Open

Read, cite, and build on anything we publish.

  • 10 peer-reviewed publications, open access
  • All 4 research themes
  • Benchmarks, datasets, prompts, and eval rigs — fully documented
  • Lab reading list (quarterly digest)

For teams

Workshop

On-site sessions for product and ML teams on evaluation, safety, and agent reliability.

  • Half-day to multi-week formats
  • Designed for product & ML teams shipping LLMs

For enterprises

Partner

Applied research collaborations with foundations, governments, and large companies.

  • Tools co-built with engineering partners
  • Foundations, governments & global corporates

Research Themes

The questions the lab is built around.

  • 01 / 04

    Benchmarks and Evaluation

    We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.

  • 02 / 04

    AI Safety and Security

    From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.

  • 03 / 04

    Agentic AI for Science

    We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.

  • 04 / 04

    Human–AI Interaction

    We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.

Research

ICML spotlights, Nature Medicine, NeurIPS Datasets & Benchmarks, ICLR, EMNLP. Click through and you're on arXiv or OpenReview, not a request form.

Search all research
all links resolve direct to the venue

Team

15 researchers + 1 PI
Prof. Adam Mahdi
Lab Lead · Associate Professor

Prof. Adam Mahdi

Oxford Internet Institute, University of Oxford

Adam leads OxRML. The group studies how language models reason, how people work with them, and how agentic systems behave on real scientific and decision-making tasks.

News

11 entries · updated quarterly
  1. 01May 2026
    Paper

    Papers accepted at ICML 2026!

  2. 02April 2026
    Conference

    OxRML at ICLR 2026

  3. 03February 2026
    Award

    Ryan Othniel Kearns Wins MSc Thesis Prize

Publications and partners.

Venues that have published our research and institutions we collaborate with.

  • 01University of OxfordHost institution
  • 02Oxford Internet InstituteAffiliated department
  • 03Nature MedicinePublished 2026
  • 04ICMLSpotlight & papers, 2026
  • 05NeurIPSDatasets & Benchmarks, 2025
  • 06ICLRAccepted, 2026
  • 07EMNLPMultiple, 2025
AI News

From the OxRML lab.

New papers, open positions, partnership opportunities, and what we have been reading.