Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling
Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona, Yen Ting Lin
Abstract
There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods can efficiently sample complex multimodal distributions, often supported by empirical results. In this work, we argue that this interpretation is fundamentally misleading. By invoking the Jordan-Kinderlehrer-Otto (JKO) scheme and Otto calculus, we establish that the canonical WGF sampling dynamics and overdamped forward diffusion share the same density evolution and therefore inherit the same metastability and slow-mixing phenomena long understood in nonequilibrium statistical physics. We analyze this family of samplers using two complementary tools -- spectral analysis and mean first-passage time (MFPT) analysis -- and show that well-separated multimodality can induce exponentially long mixing times associated with small spectral gaps and rare inter-mode transitions. For the commonly adopted log-linear annealing schedule studied here, we find that introducing intermediate distributions does not remove the exponential scaling of the total transport time. The limitation is structural rather than implementation-specific: purely local, gradient-driven transport mechanisms can require exponentially long times to transport probability mass across well-separated modes. We argue that this represents a fundamental limitation of WGF- and FODP-based sampling in their standard forms, and motivates future development of fundamentally nonlocal mechanisms for efficient multimodal sampling.
Create a lesson
Related papers
Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models
Zuokai Wen, Louis Grenioux, Weinan E et al.
The hidden advantage of mask resampling: a theory of masked autoencoders
Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová
Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions
Leyang Wang, Yakun Wang, Song Liu et al.
Optimal Transport Meets Reinforcement Learning: A Survey
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe et al.
Posterior sampling by source-space MCMC via prior-based few-step transport maps
Hoang Phuc Hau Luu, Marcelo Hartmann, Zhongjian Wang
Tolerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening
Jie Tang, Chuanlong Xie, Lixing Zhu