Building High-Performance, Cost-Efficient, and Self-Evolving AI Infrastructure by Mr Youhe Jiang
Abstract
The growing diversity and scale of AI workloads demand infrastructure that is both efficient and adaptable. This talk presents Youhe Jiang's research on aligning system execution with LLM workloads and hardware resources, while dynamically adapting execution as workloads and hardware availability change.
Through advances in distributed LLM training and serving, he demonstrates how these principles improve performance and reduce costs, with impact across open-source communities and large-scale production systems.
The talk concludes with a vision for supporting emerging AI workloads and leveraging LLMs to build self-evolving systems that continuously improve their own execution.
Biography
Youhe Jiang is a PhD student in Computer Science at the University of Cambridge, advised by Dr. Eiko Yoneki. His research focuses on machine learning systems, particularly efficient LLM serving, large-scale distributed training, and heterogeneous computing. His work has appeared at leading conferences including OSDI, EuroSys, ICML, MLSys, NeurIPS, ICLR, and VLDB.