Attention-Level Causal Intervention Framework for Multimodal Fake News Detection
Abstract
Multimodal fake news detectors may learn biased dependencies from imbalanced event distributions and incidental text–image associations, causing attention to capture dataset-specific patterns rather than reliable discriminative evidence. To address this problem, we propose an Attention-level Causal Intervention Framework (ACIM), which performs causal adjustment directly within the attention learning process. Unlike prior causal debiasing methods operating at the feature-representation level, ACIM intervenes in attention distributions where cross-modal bias emerges. By treating attention representations as mediators, ACIM applies front-door causal intervention to mitigate confounding effects and estimate attention-level intervention effects without requiring fully observed confounders. This principle is implemented through a Causal Attention Layer Module (CALM), integrated into BERT-based textual and Swin Transformer-based visual encoders to jointly model in-sample and cross-sample attention. A causal-aware fusion layer further reconstructs cross-modal attention to suppress misleading text–image co-occurrence patterns. Experiments on Twitter and PHEME achieve accuracies of 0.906 and 0.909, improving upon the strongest reported accuracy baselines by 0.9 and 0.6 percentage points, respectively, while maintaining competitive precision, recall, and F1 performance. Ablation and sensitivity analyses further support the contribution and stability of the proposed approach.
// Source
Authors: Siqi Hao, Shuohao Li, Rongxin Lin, Jun Zhang, Xianghan Wang
Institutions: National University of Defense Technology