Ridge Regression Outperforms Optimal Hard Sample-PC Selection in High-Dimensional Isotropic Gaussian Models
Abstract
We study response-guided selection of sample principal components in isotropic Gaussian linear models with dimension proportional to sample size. Under an exact-cardinality constraint, we derive the risk-minimizing rule among coordinatewise binary selectors based on each sample eigenvalue and fitted coefficient. Although selection and fitting use the same response, the realized prediction risk of the resulting procedure has a deterministic high-dimensional limit. At any prescribed retained fraction, its risk is no greater than that of top-m principal component regression (PCR), and is strictly smaller except at the endpoints and at the fraction induced by the optimal eigenvalue cutoff. When the retained fraction is also optimized, the response-dependent part of the rule disappears and the selector reduces to this cutoff. Thus response guidance improves the allocation of a fixed selection budget, but cannot eliminate the cost of zero-or-one fitting. Optimally tuned ridge regression still has strictly smaller asymptotic risk. This ordering remains valid when both procedures use the same consistent estimate of the signal-to-noise ratio, including at interpolation. We further show that a risk score unbiased for each fixed subset becomes downward biased after subset search and produces infinite limiting risk at interpolation. Simulations confirm the predicted risk ordering and identify the lower spectral edge as the source of this instability. A cross-fitted genomic example illustrates behavior outside the isotropic model.
// Source
Authors: Ruiwei Zhang, Zhaoyuan Pan
Institutions: Renmin University of China