Oxford Internet Institute · LabEst. 2022 · Oxford, United Kingdom

Advancing the science of AI evaluation, safety, and reasoning — to improve how machines and humans decide.

Mission

Led by Prof. Adam Mahdi, our lab advances the science of AI evaluation, benchmarking, safety and security. Through rigorous empirical research, we study how LLMs and agentic systems reason, interact with humans and drive scientific discovery.

Founded
2022
Faculty lead
Prof. Adam Mahdi
Affiliation
Oxford Internet Institute
§ 01Research

Three lines, asking whether we can trust what these systems do next.

01Evaluation

Benchmarks and Evaluation

We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.

Explore
02Safety

AI Safety and Security

From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.

Explore
03Agentic

Agentic AI for Science

We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.

Explore
Flagship · Annual report

The Evaluation Index, 2026.

A field-wide audit of how today's foundation models reason under pressure — across 230+ tasks, 47 model families, and a dozen reasoning failure modes we can now reliably reproduce.

2026
230+
evaluation tasks
47
model families
12
reasoning failure modes
§ 03Engage

Four ways to work with the lab.

01

Researchers & postdocs

Collaborate on evaluation methodology, safety benchmarks, and agentic-science platforms. We host visiting researchers and run joint projects with peer labs across Europe, North America, and Asia.

Visit or collaborate
02

Industry partners

Workshops, co-built evaluation harnesses, and multi-year research partnerships for teams shipping LLM products where reliability matters — healthcare, finance, retail, and government.

Engage with OxRML
03

Policymakers

Evidence-based briefings on evaluation, model risk, and agentic deployment. We translate empirical findings into the questions regulators are actually asking.

See our policy work
04

MSc Social Data Science

A one-year master’s at the Oxford Internet Institute, combining the social sciences with computational methods. Applications are typically due in early January for September entry.

MSc programme
05

DPhil Social Data Science

Apply for fully-funded DPhil studentships at the Oxford Internet Institute. We supervise across evaluation, AI safety, agentic systems, and human–AI interaction. Applications are typically due in early January for September entry.

DPhil programme
§ 04Latest at OxRML

What we've been shipping.

ConferenceApril 2026

OxRML at ICLR 2026

AwardFebruary 2026

Ryan Othniel Kearns Wins MSc Thesis Prize

ResearchFebruary 2026

New Paper in Nature Medicine!

Stay up to date

A quarterly note from the lab. Nothing else.

New papers, open positions, partnership opportunities, and what we've been reading.

Unsubscribe in one click. We never share your email.