The Mathematics of Factors: Risk, Alpha, Detection, and Capacity
Abstract
Quantitative finance uses "factor" for two distinct objects: a risk direction along which returns covary and an alpha direction along which returns are predictable. We assign each direction gauge-invariant common-covariance strength and predictive content, separating covariance recovery from mean prediction and precisely mapping their implementation interactions. The framework joins factor-space identification, beta pricing, spectral recovery, score construction, neutralisation, breadth, transaction costs, capacity, and multiplicity control. The central results price covariance-estimation error in information-ratio units. A hinge theorem identifies risk-model transfer with a covariance-geometric cosine. In a one-spike model this coefficient is obtained in closed form from spike strength θ, aspect ratio c = N/T, and forecast alignment ρ. Transfer is non-monotone in θ: loss vanishes at both extremes and is maximised at an interior θ* at or above the BBP threshold. The most damaging directions are therefore marginally visible, not invisible. Optimal eigenvalue calibration is neither the sample plug-in nor the true spike and varies with forecast alignment, so one covariance matrix cannot be optimal for every mandate. Under separated orthogonal spikes, transfer decomposes directionwise. A relabelling theorem characterises when a traded score admits a beta-pricing representation; an aspect-ratio trilogy links spectral detection, unrestricted in-sample search, and covariance-inversion instability; and exact effective breadth is z′R⁻¹z, not a name count. Simulations verify the identities and asymptotic predictions. Two CRSP calibrations then test their scope. In the N = 600, c = 0.477 design, eigenvector overlaps differ from spherical-spike predictions and cross-direction leakage defeats additive per-factor attribution. Because the validation covariance is nearly singular and the proxy window is itself high-dimensional, the transfer shortfall is not identified as either eigenpair or specific-risk error; subspace-level transfer is defensible. This negative result motivates a generalised-spike diagnostic whose separation threshold depends on the residual spectral distribution rather than on √c. The exercises are falsification diagnostics, not market estimates. A Lean 4 project provides finite algebraic declarations; asymptotic probability, correspondence with the LaTeX statements, and empirical assumptions remain outside its certificate boundary.
// Source
Authors: Miquel Noguer Alonso
Institutions: Allen Institute for Artificial Intelligence