Applying Policy Iteration for Training Recurrent Neural Networks
I. Szita, A. Lorincz
Abstract
Recurrent neural networks are often used for learning time-series data. Based on a few assumptions we model this learning task as a minimization problem of a nonlinear least-squares cost function. The special structure of the cost function allows us to build a connection to reinforcement learning. We exploit this connection and derive a convergent, policy iteration-based algorithm. Furthermore, we argue that RNN training can be fit naturally into the reinforcement learning framework.
Create a lesson
Related papers
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Mingxuan Zhang, Xiaowen Wang, Anupma Sharan et al.
Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure
Zofia Smoleń
Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.
Ownership in AI-Assisted Everyday Tasks
Megan Wei, Melanie Subbiah, Audrey Lee et al.
PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations
Julian Eggert
Limits of Confidence in Diffusion
Russ Webb, Amitis Shidani, Alice Bizeul et al.