AI & Computingarticle2026-08-14

Evidence driven dual granularity cross modal reasoning for cultural heritage visual language understanding

Open access0 citations

Abstract

This paper addresses a critical limitation in vision-language models for cultural heritage understanding, namely the lack of explicit evidence-centered reasoning mechanisms. Existing approaches typically treat visual evidence as a post hoc explanation rather than an intrinsic component of the reasoning process, resulting in unstable reasoning paths and inconsistent answer-evidence alignment. To overcome this issue, we propose an evidence-driven dual-granularity cross-modal reasoning framework that integrates local evidence modeling and global semantic decision-making in a unified structure. Specifically, the framework introduces region-level evidence alignment through cross-modal attention and performs evidence-gated aggregation to construct a global representation guided by evidence quality. Furthermore, an answer-evidence joint reasoning head is designed to enforce structural consistency between semantic prediction and evidence localization. To enhance robustness, a perturbation consistency objective is incorporated to improve stability under linguistic, visual, and modality-level perturbations. Experiments on MUSEUM-65 and the Seeing Culture Benchmark demonstrate that the proposed method significantly improves evidence localization performance (over 12% IoU gain) and achieves superior joint accuracy compared to grounding-based baselines, while maintaining competitive semantic prediction performance with substantially lower computational cost than large multimodal models. These results validate the effectiveness of embedding evidence as an intrinsic constraint within cross-modal reasoning processes.

// Source

View paper (DOI)Open access versionOpenAlexDiscover Artificial IntelligencePublished 2026-08-14

Authors: Xue Bai, Jiahui Zhou

Institutions: University of Electronic Science and Technology of China, Chengdu University of Technology, Tianjin University of Technology and Education