Society & Economicspreprint2026-08-17

Human Evaluation as Distributed Risk Sensing: Detection Strengths and Structural Blind Spots in AI Oversight

Open access0 citations

Abstract

This paper examines human evaluation of large language models as a form of distributed risk sensing rather than comprehensive assurance. It argues that large-scale human review has an uneven detection profile: reviewers are generally effective at identifying visible, repeated, and policy-defined failures but may be less sensitive to rare, structurally subtle, and cognitively demanding failures. The paper examines four factors that shape this detection profile: finite reviewer attention, calibration variance, rarity attenuation, and the difficulty of identifying internally coherent reasoning built on flawed premises. It argues that these factors can create systematic differences between the failures that are easy to detect and those that require deeper or more sustained evaluation. A constructed AI model-training scenario is used to illustrate how an assumption-level reasoning failure can appear technically coherent, escape routine review, and potentially become part of a positive training signal. The example is illustrative and is not presented as a documented industry incident or empirical case study. The paper proposes treating human evaluation as one layer of AI risk infrastructure rather than as comprehensive assurance. It recommends explicitly accounting for detection uncertainty, supplementing routine review with stress testing and novel prompt structures, and recognizing the limits of throughput-based evaluation. This is a practice-informed conceptual analysis based on observations from structured AI evaluation workflows. It does not claim to provide statistical measurement of reviewer detection rates or empirical validation of the proposed framework.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-17

Authors: Sri Vaishnavi Akkaraju