Diff-NekRS: A Scalable Differentiable Framework for Multi-Timestep Solver-in-the-Loop Training
Junoh Jung, Riccardo Balin, Bethany Lusch, Emil Constantinescu
Abstract
Hybrid physics-machine-learning solvers improve under-resolved simulations by embedding trainable corrections into the time integration. During autoregressive inference, repeated solver-model interactions can amplify small errors, motivating multi-timestep solver-in-the-loop training. However, production solvers rarely expose the derivatives needed to backpropagate through such rollouts. We introduce Diff-NekRS, a scalable differentiable framework that embeds neural corrections directly in the GPU-accelerated NekRS incompressible-flow solver. NekRS computes the authoritative forward trajectory, a manually implemented exact discrete adjoint differentiates the supported fully discrete timestep, and LibTorch supplies neural vector-Jacobian products and parameter gradients. End-to-end Taylor and centered finite-difference tests verify the assembled gradient for two-dimensional cylinder flow (2Dcyl) and the three-dimensional Taylor-Green vortex (3DTGV) across five horizons and 12-1,020 MPI ranks. At 1,020 ranks, optimizer-enabled post-setup training updates retain 54.5%-78.0% and 80.7%-81.9% weak-scaling efficiency for 2Dcyl and 3DTGV, respectively. In 200-step autoregressive inference, the M = 50 model reduces the three-seed median terminal relative L2 velocity error by 59.2% for 2Dcyl and 12.1% for 3DTGV relative to the uncorrected coarse-grid P = 2 baseline, and retains wall-clock speedups of 5.38x and 2.49x, respectively, relative to the corresponding P = 7 configurations for equal simulated-time intervals. These results establish a verified and scalable path for multi-timestep solver-in-the-loop training that improves coarse-grid trajectory accuracy while retaining a speed advantage over the high-order reference
Create a lesson
Related papers
VNS Tokamak for Medical Isotope Production
Christopher Ehrich, Pavel Pereslavtsev, Christian Bachmann et al.
A subcell-refined entropy-residual-driven limiting strategy for high-order discontinuous Galerkin methods
Geng Liang, Rui Wang, Junjie Wang et al.
MadVfold: accelerating NLO event generation and reducing negative weights with SIMD vectorization and GPUs
Andrea Valassi
Exact ballistic energy transport and emergent XXZ dynamics in an integrable three-state chain
Hanbing Liang, Fujun Liu
TubeLab: Interactive inverse design of wind instrument bores with hard spectral constraints
Carlo Andrea Rozzi, Andrea Ferroni
Stochastic consensus dynamics for decentralized decision systems
André L. M. Vilela, Caio B. L. Silva, Kenric P. Nelson et al.