When Is Inaction a Mistake? Continuation-Aware Auditing of PPO Trading Policies
Xingfei Zeng, Xin Zhong, Nanting Li, Ziyang Zhong, Lei Xiao, Guanghui Lu
Abstract
An optimal reference may recommend trading when a learned policy chooses inaction, but the recommendation depends on information and future decisions. We introduce a four-stage audit for frozen proximal policy optimization policies without retraining. It examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives. In controlled linear-Gaussian simulations, information matching explains part of the disagreement, while continuation changes its interpretation. At unit observation noise, incumbent continuation reverses 99.3% of projected-hard missed-advantage mass; repeated projected-rule deployment improves all 50 policies. These comparisons distinguish isolated action changes from policy replacement. Historical Bitcoin/Tether (BTCUSDT) replay applies this deployment perspective to a hand-specified intervention selected using 2024 data and frozen for 2025. Daily net reward improves by 135.03 basis points, with gains in 46 of 50 policies, primarily through lower turnover costs. The audit clarifies what oracle-flagged inaction implies for deployed decision making.
Create a lesson
Related papers
Extreme-Scale Linear-Scaling Kohn-Sham DFT at 100 Million Atoms: Bridging Quantum Simulations and Experiments
Qimen Xu, Yu Zhang, Dixing Ni et al.
NFT-Based Reward Mechanisms: Sybil Farming, Vesting, and Stochastic Verification
Marco Alberto Javarone, Stefanos Leonardos, Carmine Ventre
Transducer Placement and the Limits of a Four-State Reduced Model in Post-Flutter Piezoelectric Energy Harvesting from a Pitch-Plunge-Flap Aerofoil
Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken
TailWeather: from tail to extremes, a global climatological dataset for machine-learning weather forecasting
Zhi-Song Liu, Michael Boy, Risto Makkonen
Node-Shift-Encoding Genetic Algorithm with fuzzy-enhanced reference tour to solve the bi-objective service-oriented TSP
Souad Abdoune, Menouar Boulif
A Framework for Discharge Time Prediction of Energy Storage Units Based on Coupled Dynamics and Multi-Factor Aging Models
Jiaye Yang, Hansheng Su, Wangzi Zhu