The science of AI systems that earn trust.
Led by Prof. Adam Mahdi, our lab advances the science of AI evaluation, benchmarking, safety and security. Through rigorous empirical research, we study how LLMs and agentic systems reason, interact with humans and drive scientific discovery.
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
A benchmark that tells real navigation apart from stochastic search when agents work over document collections.

What the lab is publishing now.
Papers accepted at ICML 2026!
Three OxRML papers accepted at ICML 2026, with one selected for a Spotlight presentation in the main track.
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
LLM self-explanations are usually dismissed as unreliable. Measured the right way, they predict model behavior.
Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
A preregistered randomized study in Nature Medicine on how reliably LLMs serve as medical assistants for the general public.
Four standing public commitments.
OxRML's research is organised around four long-running pillars. They guide what we publish, what we open-source, and what we will and won't take money to do. We update them publicly when the science demands it.
Benchmarks and Evaluation
We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
What this commits us toAI Safety and Security
From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.
What this commits us toAgentic AI for Science
We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.
What this commits us toHuman–AI Interaction
We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.
What this commits us to
“Adam leads OxRML. The group studies how language models reason, how people work with them, and how agentic systems behave on real scientific and decision-making tasks.”
15 researchers, one programme.














Three ways to work with us.
We work with industry, foundations, and government on the questions our research touches. Every engagement keeps the science we publish honest.
hello@oxrml.comWorkshops for industry teams
On-site sessions for product and ML teams on evaluation, safety, and agent reliability.
Half-day to multi-week formats. For teams shipping LLM products in healthcare, finance, retail, and government.
Book a workshopTools co-built with engineering partners
We work with engineering partners to turn lab work into tools other teams can run.
Evaluation harnesses, safety dashboards, agentic-research platforms. We build them with partners we trust, carrying the research methods through to the code.
See our buildsResearch partnerships
Applied research collaborations with foundations, governments, and large companies.
Multi-year programmes: shared roadmaps, sponsored DPhil studentships, named labs.
Start a conversation