Biologyarticle2026-08-15

Multimodal learning attention and cognitive load assessment based on brain-computer interface and AI

Open access0 citations

Abstract

Accurately perceiving learners’ psychological and cognitive states is a crucial prerequisite for personalized intelligent education. Addressing the issue that single physiological modal signals are susceptible to environmental noise and lack robustness in online learning environments, this paper aims to construct a multimodal perception framework capable of simultaneously and accurately assessing learning attention and cognitive load. This study proposes a multimodal assessment framework: a cognitive state perception model based on AI-driven physiological (EEG) and behavioral (visual gaze) dual-modal fusion. Relational Graph Convolutional Networks (R-GCN) are used to perform non-Euclidean geometric modeling of the spatial topological relationships between EEG electrodes, achieving automatic deep representation of neurophysiological features. An Artificial Intelligence-Arbitration-based Cross-Modal Attention Network (ACM-Net) is designed, which dynamically suppresses gain and intelligently weights EEG spatial features by encoding visual gaze dynamics in real time. A joint loss function is designed based on a multi-task learning architecture to achieve independent quantitative outputs of attention level and cognitive load dimensions. The model is validated using synchronous data from participants through a dual-task induction paradigm. According to the experimental results, the accuracy of each of the functionalities are 91.8% for recognizing attention and 89.4% for assessing cognitive load. Validation on public multimodal affective computing benchmarks MAHNOB-HCI, DEAP, and SEED yielded accuracy rates of 87.6, 85.2, and 88.1% respectively for attention recognition, and 84.3, 82.7, and 85.9% for cognitive load assessment, with RMSE values of 0.251, 0.273, and 0.244 across the three datasets. When compared to the best single-modal model, the RMSE value is reduced to 0.218, which represents an approx. 30.1% performance improvement over the RMSE from that baseline individual model. The ablation studies verify the most significant contributions to performance are R-GCN spatial modeling and the AI-arbitrated cross-modal attention network. Additionally, cross-individual validation indicates that the model demonstrates good generalization performance and robustness. The findings demonstrate that the unidirectional arbitration design of ACM-Net and the heterogeneous relation modeling of R-GCN produce statistically significant improvements over symmetric cross-attention architectures and homogeneous graph networks, with RMSE reduced from 0.267 to 0.218 and cross-subject standard deviations maintained below 0.041.

// Source

View paper (DOI)Open access versionOpenAlexDiscover Artificial IntelligencePublished 2026-08-15

Authors: Yabo Yang

Institutions: University of Arts