Society & Economicspreprint2026-09-08

Where the Benchmark Can Err: Human Deviations from Model Forecasts and Their Return

Open access0 citations

Abstract

Second study in a two-domain line on human-AI complementarity, following "Sharper, Not Safer" (chess, doi:10.5281/zenodo.22267000). It asks whether human deviation from a model forecast earns a return in a domain where the model itself can err, using the one ForecastBench round with individual-level human forecasts (July 2024): 40 superforecasters and 500 public forecasters on 578 resolved targets across 162 questions, against 34 language-model variants in a matched information condition. Five hypotheses were locked and pushed to a public repository before any analysis. Main results: the median forecast of each human group carries outcome information the best single model does not encompass (H2); per-forecaster gain over the model is a stable attribute within the round (H5, split-half r = 0.78 superforecasters, 0.64 public). The gain does not rise with cross-model disagreement against the single best model (H3), and whether extremizing relative to the model pays depends on the scoring rule (H4). A five-question table sets the forecasting results beside the chess results. Three amendments and one post-results diagnostic are disclosed in full. Data: ForecastBench (Forecasting Research Institute), CC BY-SA 4.0. This paper is licensed CC BY 4.0; any derived data tables released from this study inherit CC BY-SA 4.0 from the source data. Code, specification with amendments, and reports: https://doi.org/10.5281/zenodo.22600832 (GitHub: https://github.com/aaronsun923/forecast-complementarity)

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-09-08

Authors: Aaron Sun

Institutions: University of the Republic of San Marino