Explainable image aesthetic assessment via semantic segmentation-guided cross-attention fusion
Abstract
Most existing image aesthetic assessment methods rely on global representations or simple aggregation of local features. This makes it difficult to model semantic region information and provide region-level explanations. To address this issue, we propose a Semantic Segmentation-Guided Cross-Attention Fusion (SSG-CAF) method for explainable image aesthetic assessment. The method first uses SegFormer-B0 to generate semantic masks for eight semantic categories, thereby introducing explicit semantic and spatial priors. It then combines region-level features with global visual context. A cross-attention mechanism is used to model dependencies between semantic regions and the global representation. This design supports the representation of high-level aesthetic cues, including compositional balance, subject prominence and regional harmony. SHAP is further used for region-level attribution analysis. Experiments on the cleaned AVA test set showed that SSG-CAF achieved an MSE of 0.3412 ± 0.0150, an MAE of 0.4565 ± 0.0103, a PLCC of 0.6517 ± 0.0029 and an SRCC of 0.6370 ± 0.0034. Compared with the selected baselines, SSG-CAF showed stronger correlations with human aesthetic ratings under the cleaned AVA protocol. It also produced post-hoc, quantitative region-level attribution evidence that describes model behaviour.
// Source
Authors: Hao Fang, Chengcai Cao, Zixi Huang, Minjian Hong, Qing Xiong
Institutions: Wuhan Institute of Technology