Understanding Model Behaviour by Dr Fazl Barez

20 May 2026 03.00 PM - 04.00 PM South Spine LT 22 (SS2-B2-05, near LKC-LT) Current Students, Industry/Academic Partners

Abstract

AI systems are making consequential decisions, yet we lack reliable methods to verify why. This talk examines what it actually means to understand a neural network — not just describe its activations, but trace the internal cause of a specific output, intervene to fix it, and confirm the fix holds. I’ll argue that the three dominant approaches to model explanation (chain-of-thought, mechanistic reverse-engineering, and learned explainers) each capture something real but none closes the loop from observation to verified correction. Drawing on work from my group at Oxford — covering circuit discovery, concept erasure, CoT faithfulness, and safety monitoring — I’ll present a scientific methodology for model understanding.

 

About the Speaker

Fazl Barez is a Senior Researcher and Principle Investigator at the University of Oxford, where he leads the Technical Safety & Governance Lab and serves as Technical Director of the AI Governance Initiative. His research spans mechanistic interpretability, AI safety, and technical governance — focused on understanding the internal workings of neural networks and using that understanding to make AI systems more reliable and auditable. He teaches Oxford’s AI Safety and Alignment course and is Principal Scientist at Martian. His work is supported by OpenAI, Anthropic, Schmidt Sciences, and NVIDIA.