AI Safety Through Mechanistic Interpretability: Monitoring, Interpreting, and Controlling Internal Mechanisms in Large Language Models by Serino Antonio
Abstract
Ensuring that Large Language Models behave safely and align with human intentions remains a fundamental challenge, often obscured by the black-box nature of deep neural networks. In this seminar, we explore how mechanistic interpretability and representation engineering provide principled tools to monitor, interpret, and control internal mechanisms directly within transformer representations. We examine how safety features emerge along distinct geometric structures inside activation spaces. Building on these latent structures, we discuss diagnostic and interpretability frameworks, ranging from linear probes to Sparse Autoencoders, designed to isolate and track features that foster trustworthy AI while directly addressing the challenge of concept faithfulness. Finally, we offer key insights into what these geometries reveal about model controllability: contrasting mere representational decodability with true causal steering, exploring the role of localized internal components, and discussing practical implications for white-box monitoring and alignment without extensive fine-tuning.
About the SpeakerAntonio Serino is a Ph.D. student in Data Science and Artificial Intelligence at the University of Milano-Bicocca and a Visiting Researcher at Nanyang Technological University. He is an Ermenegildo Zegna Scholar and actively contributes to high-performance computing research on the Leonardo supercomputer, serving both as Principal Investigator and team member.
His research lies at the intersection of Mechanistic Interpretability, Representation Engineering, and AI Safety, investigating how Large Language Models encode safety-critical features into latent geometric spaces, how these can be faithfully interpreted, and how such geometric structures can be leveraged to steer and control internal safety mechanisms.
Despite being an early-stage researcher, he has contributed to several papers in top-tier AI and NLP venues, including EMNLP, EACL, IJCAI, and ECML-PKDD.