Differentiable Hybrid-Action Neural Feedback Control for District Heating Networks
Nicolas Kirsch, Corrado Sgadari, Alessio La Bella, Giancarlo Ferrari-Trecate
Abstract
Many cyber-physical systems require control policies that combine continuous setpoints with discrete operational de- cisions, such as equipment switching, mode selection, or resource scheduling. Discrete actions are not differentiable, which ob- structs gradient-based policy training, while conventional mixed- integer formulations remain costly to solve online. This paper proposes a hybrid-action neural controller (HANC) in which a continuous branch, a categorical branch and a differentiable assembly layer jointly generate commands that satisfy complex actuator constraints by construction. Categorical decisions are handled using a straight-through Gumbel estimator, enabling the policy to be trained by backpropagation through time over full closed-loop rollouts. The proposed framework is deployed on a district heating network (DHN) featuring multiple heat generation units and stratified thermal energy storage. Its performance is evaluated on a simulation of a real DHN located at RSE SpA in Italy. The resulting policy jointly learns switching decisions and continu- ous operating setpoints. Under dynamic electricity pricing, the learned controller reduces operating cost by 30% compared to a rule-based industrial baseline. We also show that, compared with a deterministic straight-through relaxation, injecting noise during training achieves similar cost while reducing hard switching by an order of magnitude, and attribute this difference to the wider decision margins of the resulting policy.
Create a lesson
Related papers
Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
Alexey Peregudin, Ngoc Tuan Dinh
Feasibility of Simultaneous Input-Output Constraints for Tracking in a Class of LTI Systems: Part I
Junhyeok Yoon, Heather Hussain, Anuradha M. Annaswamy
Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation
Yuta Takahashi, Shin-ichiro Sakai
Local Consistency Does Not Guarantee Global Conservation: Auditing Zero-Shot Composition of Airway Flow Operators
Nichula Sathmith Wasalathilaka, Navodya Heshan Samarasinghe Dhanujaya Suraweera, Kevin Dawson et al.
A Dynamic Generalized Kalman Consensus Filter for Switching Sensor Networks
Tirthankar Chakraborty, Arnab Maity
Interactive Power Flow in the Browser
Samuel Talkington, Frederik Geth, Qian Zhang et al.