Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs From Human Behavior
Pranjal Rawat, John Rust
Abstract
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.
Create a lesson
Related papers
Shrinkage Bayesian Causal Forest with Instrumental Variable
Lennard Maßmann, Jens Klenke
Conditionally linear, matrix normal state space models
Drew D. Creal, Marcelo C. Medeiros, Rodrigo Sarlo
Policy Targeting with Market Equilibrium
Gyungbae Park
What No First Stage Can Detect: Functional-Form Contamination in Linear IV
Parush Arora
Tensor-BEKK: Conditional Covariance Modeling and Inference for Tensor-Valued Time Series
Huan Gong, Feiyu Jiang
Profiled Anderson--Rubin Test: Robust Inference Allowing for Direct Effects of Instruments
Jung Hyub Lee