A hybrid deep learning approach for interpretable inference of potential accidents from hazard reports in hydropower engineering construction
Abstract
Hazard reports provide early evidence of unsafe conditions and behaviours in hydropower construction, but their potential for proactive accident prevention is often constrained by unstructured narratives and reliance on manual interpretation. Using 9,606 Chinese-language hazard reports collected from the Baihetan Hydropower Station as a case study, this study develops an interpretable hybrid deep learning approach for inferring 15 potential accident categories. The proposed model integrates RoBERTa for contextual semantic representation, a bidirectional long short-term memory network for sequential dependency modelling and hierarchical attention for multi-level feature aggregation and multi-label classification. Experimental comparisons with baseline models show that the proposed approach achieves mean precision, recall and F1-score values of 89.83%, 88.12% and 88.97%, respectively, exceeding the strongest baseline, RoBERTa + BiLSTM, by 1.91 percentage points in F1-score. Ablation experiments further demonstrate the complementary contributions of contextual semantic representation, sequential dependency modelling and attention-based feature aggregation. SHapley Additive exPlanations (SHAP) are further employed to identify influential words and phrases contributing to model predictions, thereby improving prediction transparency and interpretability. The findings demonstrate the potential of transforming unstructured hazard records into actionable safety information to support risk monitoring, targeted intervention and evidence-based safety management in large-scale hydropower construction projects.
// Source
Authors: Jiayi Zhou, Zhiyi Ge
Institutions: Huazhong University of Science and Technology, Tianjin University