Physics & Spacepreprint2026-08-23

Prediction-Horizon Transport of Adam's Optimal Momentum State

Open access0 citations

Abstract

Adam's future behavior depends not only on the model parameters but also on its hidden first-moment state. This work develops a prediction-horizon transport theory for Adam's finite-horizon optimal momentum state under frozen local quadratic dynamics in a prescribed correction subspace. We derive an exact geometry for the horizon-indexed optimal state and show that the excess frozen loss produced by using an approximate target-horizon state is exactly one half of its squared distance from the target optimum in the corresponding quadratic metric. In the low-curvature limit, the optimal absolute momentum state evolves radially across prediction horizons according to an explicit temporal coefficient determined by Adam's bias-corrected first-moment dynamics. The coefficient depends on the prediction horizons, training step, and beta1. More generally, learning rate and scalar effective curvature enter the full horizon response through their product. Finite curvature and subspace leakage generate departures from radial transport, while additional block-Krylov information provides systematic finite-curvature refinement. On four held-out Burgers physics-informed neural-network checkpoints, transporting an H = 5 restricted optimal momentum state to H = 25 produced only 0.015% to 0.605% normalized squared state mismatch in the target-horizon quadratic metric relative to the independently computed H = 25 state. With deployment fixed at half the computed correction and no trust projection, the transported state reproduced approximately 99% of the improvement in the nonlinear 50-step objective obtained by the exact H = 25 state, while requiring 16 rather than 96 Hessian-vector products in the current implementation. Separate training trajectories at beta1 = 0.8, 0.9, and 0.95 followed the parameter-free predicted change in the radial transport coefficient, with increasing finite-curvature deviation at beta1 = 0.95. The results establish a cheap longer-horizon state predictor rather than a universally superior Adam controller. A reliable inexpensive criterion for deciding in advance when finite-curvature refinement is necessary remains an open problem.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-23

Authors: Caner Sakar