Diagnosing Motor-Imagery BCI Failure: A Three-Step Diagnostic Framework
Abstract
Version 3.1.0. This version adds the four figures. The v3.0.0 PDF carried the captions but not the artwork, because the manuscript is prepared with figures supplied as separate files, which is what the journal requires and what a reader does not want. The text is otherwise the version submitted to the Journal of Neuroscience Methods on 2 August 2026, with one sentence added in Section 2.4 so that Figure 1 is cited before Figure 2 and the figures are numbered in order of first citation. Versions from 3.0.0 onward describe a three-step framework evaluated on three public cohorts (75 subjects); v2.0.0 and earlier described a two-step framework on a single cohort. This is a preprint and has not been peer reviewed. Background. A substantial minority of users cannot achieve reliable control with motor-imagery BCIs. Existing approaches predict or treat failure, but none asks whether failure is attributable to the decoder or to the signal itself before intervention. New Method. I propose a three-step procedure. Step 1 computes classDis (Lotte and Jeunet 2018) from training-session covariance matrices. Step 2 adds a transfer gate comparing within-session and cross-session CSP-LDA performance. Step 3 applies paired permutation tests for FBCSP and separate tests for MDM, with multiple comparison correction on gate-passing subjects. Results. Pooled across three cohorts (BCI IV 2a, BNCI2015-001, Lee2019; 75 subjects), classDis correlates with cross-session decoding at r = .66, 95% CI [.50, .77]. In Lee2019, 11 subjects are flagged (20.4%), of whom four are recovered: three by FBCSP (S12, S17, S20) and one by MDM (S5), none by both. The remaining seven are transfer failures. BNCI2015-001 has one flagged subject (S11), recovered by neither decoder. BCI IV 2a has no flagged subjects. Comparison with Existing Methods. Against Blankertz et al. (2010), classDis predicts cross-session decoding rather than online feedback performance. Against Vidaurre and Blankertz (2010), classDis is classifier-free and measured before any decoder is trained. Against Ang et al. (2012), their group-level FBCSP-vs-CSP test on 2a evaluation data was non-significant at p = .059, whereas S5 is p < .001 in this framework's per-subject test. Conclusions. Failure is heterogeneous and the distinction is measurable. Any diagnostic label is relative to a specific recording and pipeline, never a claim about the person. The analysis code and every derived result table are archived separately at https://doi.org/10.5281/zenodo.21760836. The EEG data are obtained through MOABB and are not redistributed.
// Source
Authors: Peilin Zhong
Institutions: Shadyside Hospital