Simulation-Based Neural Policies for Portfolio Choice: Architecture, Training, and Interpretability
Jules Viard, Alexander Michaelides, Panos Parpas
Abstract
Many economic decision problems, lifecycle consumption-saving and dynamic portfolio choice, are finite-horizon stochastic control problems with continuous states and actions. When the state is low-dimensional these problems are solved by dynamic programming on a grid. The grid cost grows exponentially in the state dimension, known as the curse of dimensionality, which motivates replacing the value-function grid with a neural policy optimized directly through simulation. Such policies are usually studied in the high-dimensional settings that motivate them, precisely where no reference solution exists. So the contribution of any single architectural or training choice cannot be isolated and diagnosed. We therefore take a step back and treat both the architecture and the solution method as the objects of study. To this end, we consider a lifecycle problem with a sufficiently low-dimensional normalized state space to admit an accurate dynamic programming solution, which is used for evaluation. We compare four architectures. The simplest consists of a single time-conditioned network. We then consider two networks concatenated across the regime switch, followed by one network per date trained backward against frozen downstream policies. Finally, we evaluate a constrained variant of the per-date architecture. Decoupling the policy across time gives each date a short, well-posed objective, which we pair with direction-dominant optimization that normalizes away gradient magnitude. Architectures that lead to similar realized utility objective can nevertheless differ in whether they respect the underlying problem's economics. We therefore evaluate each design jointly based on welfare, a solution-free Bellman residual, shape restrictions, and the resulting policy functions.
Create a lesson
Related papers
Level-Set Geometry and the Theoretical Performance of PDHG for Conic Linear Optimization
Zikai Xiong, Robert M. Freund
Trajectory Manifolds for Nonlinear Data-Enabled Predictive Control
Arda Bayer
Optimizing Lyapunov Certificates via Stability-Preserving Quadratization for Polynomial Systems
Yubo Cai, Gioele Zardini
Regularity of a Multidimensional Principal-Agent Problem with Separable Effort Costs
Shuaijie Qian, Guan Qiao
Near-Optimal Exact-Value Zeroth-Order Complexity for Smooth Strongly Convex Optimization
Wendao Wu, Haihan Zhang, Chenheng Zhang et al.
A VU-calculus for composite functions and the U-Hessian of partly smooth functions
Shuai Liu