Benchmarks and Evaluation
We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
- Benchmark design
- Statistical evaluation
- Capability elicitation
- Contamination audits
@ University of Oxford
Led by Prof. Adam Mahdi, we work on the science of evaluating, benchmarking and securing modern AI. Our empirical research asks how LLMs and agentic systems reason, collaborate with humans and accelerate scientific discovery — alongside industry partners deploying these systems where reliability matters.
We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.
We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.
We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.
Two ways to work with us: third-party evaluation of your models and agents, or a focused engagement that turns one of our research outputs into a tool you own.
We accept students through Oxford Internet Institute graduate programmes. Applications are typically due in early January for September entry.
A one-year master’s at the Oxford Internet Institute, combining the social sciences with computational methods.
A fully-funded doctorate. Supervision spans evaluation, AI safety, agentic systems, and human–AI interaction.