LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes
Youcheng Zong, Runda Jia, Dakuo He
Abstract
Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from limited interactions which process variables each action affects, in which direction, and after what delay. Fixed industrial documents already describe part of these relations, but their open-text statements neither represent the current operating condition nor directly fit a numerical policy. This article presents LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes (LCAE), which uses a large language model before training to normalize fixed documents into a frozen action--observation--direction--delay relation basis. Recent numerical action--response history then modulates the current strength of each relation, while the evaluated action forms a state-conditioned nonlinear action-effect field in the same basis. The critic evaluates actions through this field, and the actor uses the same relation gains to generate actions, making document semantics part of maximum-entropy policy learning. Neither the LLM nor the embedding model runs online during training or deployment; the deployed policy uses only frozen semantic artifacts and visible numerical history. The method states a falsifiable hypothesis: when documented relations are correct and recent history reflects their contextual strength, this action representation should provide a more useful decision bias than raw action coordinates.
Create a lesson
Related papers
Leader-Follower Formation Control with Prescribed Convergence Rates under Bearing Persistence of Excitation
Tarek Bouazza, Zhiqi Tang, Soulaimane Berkane et al.
On asymptotic stability of the time-varying Kalman filter for unstabilizable linear systems: an optimization perspective
James B. Rawlings, Titus Quah, Matthias A. Müller
Designing Grid-Aware Dynamic Specifications for Large Data Center Loads
Ashutossh Gupta, Vassilis Kekatos
Time-Optimal Operation of a Load-Hoisting Gantry Crane
Eric Mountain, Tarunraj Singh
Learning to Solve Two-Stage Stochastic Unit Commitment Problems with Quality Guarantees
Andrea Fusco, Andrea Lodi, Lavanya Marla
Towards Interaction Regulation from Human Feedback via Free Energy Minimization
Maria Paula Diaz Monfort, Cinzia Tomaselli, Michael Richardson et al.