Seminar: Flexible and Scalable Reinforcement Learning Systems
Machine learning (ML) systems translate data into value for decision making. Recent breakthroughs in large ML models (e.g., GPT 4, Llama 3, Gemini) and the remarkable outcomes of reinforcement learning (eg., AlphaFold, FunSearch, AlphaGeometry) have shown that scalable and flexible MLtraining/inference on the industrial scale (e.g., tens of thousands of GPUs/accelerators) is critical to obtain state-of-the-art performance. This talk aims to answer the question “how to co-design multiple layers of the software/system stack to improve the flexibility, scalability and the performance of ML computation”. It addresses the challenges to design and build efficient ML systems that integrate the scalable ML layer, the distributed data/state management layer, and the compilation-based optimization layer.
Biography