AI & Computingarticle2026-09-02

From Alarms to Probabilities: Stratified Human Review, Label-Noise Correction, and Calibrated Risk Grading for Industrial Vibration Anomaly Detection

Open access0 citations

Abstract

Industrial anomaly detectors emit binary alarms, but production lines need graded dispositions. Calibrating alarms into fault probabilities requires ground truth, coming only from noisy human review treated as exact. We report a deployed closed loop on a reciprocating-compressor line (46,023 units, nine test campaigns, four fused detector legs). A stratified review of 1161 units audited a fusion-score grading and refuted its assumed monotonicity: precision was 25.2%/13.8%/26.1% for high/medium/low tiers. The cause was correlated false positives: two legs firing on shared broadband transients agreed on 480 units at 21.0% precision—detector agreement is not independent evidence—whereas one periodicity feature was monotone (27.8% → 60.0% → 100%). A blind test against seeded fault units (hardware ground truth) measured reviewer sensitivity at 0.905 and specificity at 0.421 on hard cases; Rogan–Gladen correction restored monotonicity (95% of bootstrap replicates; 83% under campaign-cluster resampling), exposed the low-tier advantage as a label-noise artifact, and re-estimated no-alarm prevalence at 4–10% versus the observed 13.3%. Corrected evidence drove a redeployed rule calibrating tiers at ≈67%/27%/17% under review budgets (≤1%/≤3%/≤9% of production). The methodology—stratified audit, seeded-fault blind testing, and evaluation-side prevalence correction—transfers to any human-verified system.

// Source

Authors: Tao Feng, Kun Chen, Jing Wang, Haonan Guo, Jiewen Wen, Tong Ji

Institutions: Beijing Technology and Business University