How Diffusion Models Learn Low-Dimensional Structure
Diffusion models achieve remarkable performance on high-dimensional data that often concentrate near low-dimensional structures. Yet it remains unclear how such models discover and exploit this structure during learning. In this talk, I will present recent theoretical results characterizing the score function of low-dimensional data across different noise scales. At small noise levels, the dominant component of the score points in the normal direction toward the underlying data manifold, and the corresponding denoising map approaches a manifold projection. As the noise scale increases, information about the probability density along the manifold becomes increasingly important. This reveals a collapse-and-refine mechanism: diffusion models first recover the geometric support of the data and subsequently refine the probability distribution within that structure. This multi-scale characterization provides a concrete explanation for how diffusion models can exploit low-dimensional structure despite operating in a high-dimensional ambient space and clarifies the distinct roles of geometry recovery and density learning in diffusion models.