PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
Abstract
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
Create a lesson
Related papers
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Tianjie Ju, Zheng Wu, Yueqing Sun et al.
Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information
Chanho Park, Daehyeon Choi, Jihyun Lee et al.
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
Agniv Chatterjee, Georgios Pavlakos
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Lukas Kuhn, Lucas Maes, Giuseppe Serra et al.
Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
Frederik Berenz
KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations
Chenchen Ge, Hanwen Shen, Bowen Jing et al.