For individuals
Read, cite, and build on anything we publish.
- 10 peer-reviewed publications, open access
- All 4 research themes
- Benchmarks, datasets, prompts, and eval rigs — fully documented
- Lab reading list (quarterly digest)
Led by Prof. Adam Mahdi, the Reasoning with Machines Lab advances the science of AI evaluation, benchmarking, safety and security. Through rigorous empirical research, we study how LLMs and agentic systems reason, interact with humans, and drive scientific discovery. We work with industry partners deploying AI where reliability matters.

Adam leads OxRML. The group studies how language models reason, how people work with them, and how agentic systems behave on real scientific and decision-making tasks.
Read, cite, and build on anything we publish.
On-site sessions for product and ML teams on evaluation, safety, and agent reliability.
Applied research collaborations with foundations, governments, and large companies.
The questions the lab is built around.
We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
From bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.
We build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.
We run large-scale empirical studies on how people use AI for high stakes decisions, from healthcare and law to policy and beyond.
ICML spotlights, Nature Medicine, NeurIPS Datasets & Benchmarks, ICLR, EMNLP. Click through and you're on arXiv or OpenReview, not a request form.

A benchmark that tells real navigation apart from stochastic search when agents work over document collections.
Ł Borchmann, J Van Landeghem, M Turski, S Padarha, RO Kearns, A Mahdi, et al.

LLM self-explanations are usually dismissed as unreliable. Measured the right way, they predict model behavior.
H Mayne, JS Kang, D Gould, K Ramchandran, A Mahdi, NY Siegel

A benchmark that obfuscates orthography to strip memorised knowledge out of reasoning problems, showing how much "reasoning" was recall.
J Khouja, K Korgul, S Hellsten, L Yang, V Neacsu, H Mayne, RO Kearns, A Bean, A Mahdi
Papers accepted at ICML 2026!
OxRML at ICLR 2026
Ryan Othniel Kearns Wins MSc Thesis Prize
Venues that have published our research and institutions we collaborate with.
New papers, open positions, partnership opportunities, and what we have been reading.