What Determines Noise Robustness in Symbolic Regression? A Controlled Study of Physical Law Rediscovery
Abstract
Symbolic regression promises to rediscover interpretable physical laws directly from data, but real experimental measurements are noisy, and it is unclear which properties of a governing equation make it easier or harder to recover under noise. We conduct a controlled study using PySR across 17 physical laws spanning four hypothesized complexity tiers, at 8 noise levels and 3 sample sizes (6,120 runs total), scoring both exact symbolic recovery and approximate (R² > 0.99) recovery. As expected, recovery degrades sharply past a noise ceiling (present in every equation tested), and nested, non-closed-form equations are essentially never recovered exactly (0/720 runs, Fisher's exact p = 1.8×10⁻²³³). Our initial hypothesis — that noise tolerance tracks a tier structure built from variable count and operator nesting — does not hold as a ranking: each tier's aggregate noise-tolerance coefficient changes sign or significance depending on how many constant-free equations it happens to contain, which is itself the key finding. Tracing one striking anomaly (Coulomb's law recovering only 33% of the time even on clean data) to its source, we identify whether the target expression requires fitting a non-trivial numeric constant as a specific mechanism, and test it directly: across 6 matched equation-family pairs, the constant-free member recovers exactly at every noise level tested in 5 of 6 families outright, and a paired Wilcoxon signed-rank test across all 6 families is significant (p = 0.016). A primary interaction model makes the effect size concrete: constant-bearing equations' noise-sensitivity coefficient is roughly 11 times larger than constant-free equations', translating to model-predicted recovery rates of 95.0% vs. 2.0% at 30% noise — a finding corroborated by a formal AIC model comparison, a mixed-effects robustness check, and an outlier-robust cross-equation correlation analysis (n=16 equations). Separately, noise thresholds scale substantially with sample size (Kruskal-Wallis H = 14.9, p = 0.0006), meaning claims about which equations are "noise-sensitive" are only meaningful relative to a stated data budget. Together these results indicate that constant-fitting burden, not structural complexity as originally operationalized, is the dominant driver of noise robustness in this setting.
// Source
Authors: Ashton Tovar Burke