HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC
Ming Tan, Xiyun Jiao
Abstract
Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning methods are not directly applicable. We propose HyperMC, a multi-fidelity tuning framework that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. By running multiple successive-halving brackets, HyperMC balances broad exploration of a continuous hyperparameter space with increasingly accurate evaluation of promising configurations under a fixed computational budget. We further introduce Robust HyperMC, which uses global grid initialization followed by elite-guided local refinement to reduce sensitivity to random candidate generation and noisy finite-budget evaluations. Under suitable approximation and concentration conditions for the estimated KSD, we establish that the successive-halving component selects a near-optimal configuration among the sampled candidates with high probability and derive a sufficient computational budget for successful selection. Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.
Create a lesson
Related papers
Copula Transformations for Data-Consistent Inversion
Troy Butler, Tianyi Jiang, João Silva et al.
Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing
Zhaoming Li, Paul Hand
Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
Jia-Nan Wang, Zixun Huang, Kairui Li et al.
A computational approach to maximum likelihood thresholds for colored Gaussian graphical models
Roser Homs, Olga Kuznetsova, Bernadette J. Stolz
From topology learning to graph generation: A unifying perspective
Xiaowen Dong, Hoi-To Wai, Siheng Chen et al.
Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation
Shizhe Zhang, Mingyang Zhao, Lei Ma