Benchmarks and Evaluation
We develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
ExploreWe develop the science of LLM evaluation, setting the standard for rigorous assessment and identifying hidden risks before they matter.
ExploreFrom bias and toxicity to agentic misalignment, we study the full spectrum of AI risk and develop the technical and governance tools to address it.
ExploreWe build agentic systems that automate scientific knowledge synthesis and discovery, with a focus on agents that are reliable, transparent and domain-grounded.
ExploreWe're selective by design. Each track has its own intake cadence — pick the one closest to how you work.
Collaborate on evaluation methodology, safety benchmarks, and agentic-science platforms. We host visiting researchers and run joint projects with peer labs across Europe, North America, and Asia.
Visit or collaborateWorkshops, co-built evaluation harnesses, and multi-year research partnerships for teams shipping LLM products where reliability matters — healthcare, finance, retail, and government.
Engage with OxRMLEvidence-based briefings on evaluation, model risk, and agentic deployment. We translate empirical findings into the questions regulators are actually asking.
See our policy workA one-year master’s at the Oxford Internet Institute, combining the social sciences with computational methods. Applications are typically due in early January for September entry.
MSc programmeApply for fully-funded DPhil studentships at the Oxford Internet Institute. We supervise across evaluation, AI safety, agentic systems, and human–AI interaction. Applications are typically due in early January for September entry.
DPhil programme