The Mean–Median Midpoint Estimator: Exact Risk Geometry, Sharp Class Constants, and Unexpected Gains from Equal Weighting
Abstract
This paper studies the mean–median midpoint estimator, defined as the equal-weighted average of the sample mean and sample median, as a statistical decision rule under squared-error loss. The estimator itself is classical and is not generally optimal. The purpose is to determine when this fixed midpoint is competitive, when it can outperform both constituent estimators, how large those gains can be over broad distribution classes, and what information should move the weight away from one half. The first part develops an exact risk geometry for two-estimator averaging. For any two square-integrable estimators of the same scalar target, the midpoint beats both constituents exactly when their mean-squared disagreement exceeds twice the absolute difference between their individual mean squared errors. This is a finite-sample result requiring no distributional model, asymptotics, or symmetry. At exact risk balance, any nonzero disagreement makes the midpoint the unique affine risk minimizer. In regular large-sample settings, the same geometry reduces to the imbalance between endpoint variances and the correlation of their errors, producing a decision map that quantifies both the gain from using the midpoint and its cost when another endpoint should be preferred. Under bounded risk imbalance and bounded error alignment, the midpoint also has a sharp competitive-minimax guarantee. Under symmetric location models, the mean–median problem reduces to two scale-free shape characteristics: central density, which controls the mean–median variance balance, and standardized mean absolute deviation, which gives their asymptotic correlation. At the point where the mean and median have equal asymptotic variance, the log-concave case adds exact attainment: the midpoint reduces variance relative to either endpoint by between 10.280% and 13.520%, and both bounds are attained by explicit extremal distributions—a two-sided truncated exponential at the lower end and a uniform-core/exponential-tail distribution at the upper end. At the same mean–median variance crossover, the corresponding sharp open ranges are 6.699%–25% over symmetric unimodal distributions and 10.106%–18.169% over Gaussian scale mixtures; in both classes the endpoints are approached but not attained. The attained log-concave result follows from a reduction to a moment-extremal problem and the log-concave moment theorem of Eskenazis, Nayar, and Tkocz. The symmetric log-concave class also shows directly why the admissible set of estimators matters. If the available choices are only the mean, the median, and their midpoint, the midpoint is minimax relative to the better endpoint, with sharp worst-case factor 1.75. If every fixed convex mean–median weight is allowed, the minimax rule instead places two-thirds of the weight on the mean and one-third on the median, with sharp worst-case factor 13/9 (about 1.444). Both constants are attained. The midpoint can therefore have a precise minimax justification in one decision problem without being a universally optimal weight. The paper further examines parametric efficiency across Student-t, generalized-Gaussian, logistic, hyperbolic-secant, and Gaussian-mixture location families; finite-sample behavior; comparisons with Huber, trimmed-mean, and Hodges–Lehmann estimators; fixed-weight minimaxity over specified model classes; retained-summary equivariance; and the consequences of estimating the combination weight. These analyses identify regimes in which the midpoint coincides with the affine oracle, strictly outperforms both constituents, or attains high efficiency, as well as regimes in which additional structure favors a directional weight or a different estimator. The limitations are explicit. The midpoint retains an unbounded influence function and zero asymptotic breakdown; material mean–median separation can make it inappropriate for a population-mean target; infinite variance prevents finite mean-squared-error and root-n guarantees; and sufficiently severe contamination can shift the class-minimax fixed weight substantially toward the median. The interpretation is therefore conditional rather than universal. When both classical summaries remain plausible and the available information does not justify a directional weight, the mean–median midpoint estimator is a transparent fixed rule that can deliver genuine risk reduction without an additional weight-estimation step once the mean and median are available; when reliable structure points elsewhere, another weight or estimator may be preferable. The accompanying open-science release contains the complete manuscript source, all nine publication figures and their generators, simulation programs and machine-readable outputs, and numerical verification of the principal analytical constants, regime boundaries, class-minimax calculations, and simulation-derived tables.
// Source
Authors: Ketan Singh