§01 · Research

Research

  • 2025

    LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

    H Mayne, RO Kearns, Y Yang, AM Bean, E Delaney, C Russell, A Mahdi

    September 2025EMNLP

    Ask an LLM "what would change your answer?" and it looks like introspection. It is often confabulation.

  • How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

    Y Yang, F Sondej, H Mayne, A Lee, A Mahdi

    November 2025EMNLP

    Direct Preference Optimization reduces toxicity. We trace where it acts, neuron by neuron.

  • Evaluating LLM-as-a-Judge under Multilingual, Multimodal and Multi-domain Constraints

    S Padarha, E Semenova, B Vidgen, A Mahdi, S A Hale

    November 2025NeurIPS LLM Lifecycle Workshop

    How LLM judges degrade across languages, modalities, and domains, and where the failure modes sit.

  • Measuring what matters: Construct validity in large language model benchmarks

    AM Bean, RO Kearns, A Romanou, FS Hafner, H Mayne, J Batzner, et al.

    November 2025NeurIPS Datasets and Benchmarks

    A construct-validity audit of widely-used LLM benchmarks: what they claim to measure versus what they capture.

  • 2026

    A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

    H Mayne, JS Kang, D Gould, K Ramchandran, A Mahdi, NY Siegel

    May 2026ICML

    LLM self-explanations are usually dismissed as unreliable. Measured the right way, they predict model behavior.

  • Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

    Ł Borchmann, J Van Landeghem, M Turski, S Padarha, RO Kearns, A Mahdi, et al.

    May 2026ICML (Spotlight)

    A benchmark that tells real navigation apart from stochastic search when agents work over document collections.

  • 2025

    It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

    K Korgul, Y Yang, A Drohomirecki, P Błaszczyk, W Howard, L Aichberger, C Russell, P H S Torr, A Mahdi, A Bibi

    May 2025ICML

    A benchmark for whether web agents can be socially engineered into abandoning the user's task. Today's agents fall for it.

  • 2026

    Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

    AM Bean, RE Payne, G Parsons, HR Kirk, J Ciro, R Mosquera-Gómez, S Hincapié, AS Ekanayaka, L Tarassenko, L Rocher, A Mahdi

    February 2026Nature Medicine

    A preregistered randomized study in Nature Medicine on how reliably LLMs serve as medical assistants for the general public.

  • 2025

    Review of multimodal machine learning approaches in healthcare

    F Krones, U Marikkar, G Parsons, A Szmul, A Mahdi

    February 2025Information Fusion

    A survey of multimodal ML in clinical practice, from data-fusion strategies through to deployment.

  • 2026

    LingOly-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

    J Khouja, K Korgul, S Hellsten, L Yang, V Neacsu, H Mayne, RO Kearns, A Bean, A Mahdi

    April 2026ICLR

    A benchmark that obfuscates orthography to strip memorised knowledge out of reasoning problems, showing how much "reasoning" was recall.

§ Next

Want to collaborate on a paper or dataset?