Post-Stability Representational Change]{When Accuracy Stabilizes, Representations Still Move: Learning-Rule-Dependent Geometry of Post-Stability Representational Change
Abstract
Neural networks are usually considered "converged" once validation accuracy plateaus. This preprint asks whether the internal representations have actually stopped changing too and finds that, often, they haven't. We train shallow MLPs on binary MNIST using two learning rules: standard backpropagation and a hybrid Hebbian-gradient rule. Both reach similar accuracy. But their internal geometry after that point looks very different. Backpropagation collapses hidden representations down to a near one-dimensional manifold within a few epochs, then stays frozen. The hybrid rule keeps a higher-dimensional, slower-changing representation over a much longer window. An epoch-matched control rules out the training phase as the explanation for the difference; the learning rule itself. We measure this post-stability change ("operational drift") with representational similarity analysis, linear CKA, centroid and class-axis shifts, decoder transfer, participation ratio, and Procrustes alignment. Classification accuracy from a linear decoder stays above 98% throughout, even as the geometry keeps moving and the two rules become poorly aligned. Submitted to NeurReps 2026 (Symmetry and Geometry in Neural Representations)
// Source
Authors: Elaheh Sabbaghi, Talha Nazar