An analysis of deep-learning systems found that they often identify exons using simple sequence patterns while overlooking RNA structure.
Researchers developed an approach called interpretable distillation to examine how splicing prediction systems make their decisions. The analysis found that the systems recognize exons through simple additive combinations of sequence motifs, including known elements that regulate splicing.
But the models also used genomic confounders—patterns unrelated to splicing—and did not adequately account for RNA structure. As a result, they showed poor performance on non-reference sequences and made systematic prediction errors.
How the models make predictions
The researchers used interpretable distillation to explain the prediction logic of RNA splicing models. They found that the models recognize exons through surprisingly simple combinations of sequence motifs, including known splicing regulatory elements.
The analysis also found pervasive confounders and blind spots. The models exploited genomic patterns unrelated to splicing and failed to adequately capture the effects of RNA structure. These limitations were associated with poor performance on non-reference sequences and systematic prediction errors.
Why the blind spots matter
Splicing prediction tools are used to study gene regulation and interpret how genetic variation may affect genes. The findings indicate that strong performance on genomic sequence data does not necessarily mean that a model has learned the biological features that control splicing.
Identifying irrelevant sequence clues and missing RNA structure gives researchers specific limitations to address when developing and evaluating future splicing models. The study therefore points to ways of making these tools more reliable, particularly for sequences that differ from the reference sequence.
Evidence and caveats
This is a model-analysis study based on an interpretable distillation framework applied to RNA splicing prediction systems. The abstract reports consistent findings about sequence motifs, genomic confounders, RNA structure and errors on non-reference sequences.
The abstract does not specify how many models or sequences were analyzed, which datasets were used, or how the models’ performance was quantified. It also does not establish that every splicing prediction system has the same weaknesses, so the findings’ broader applicability cannot be assessed from the provided information.
// Source
Genome biology · 2026 · DOI: 10.1186/s13059-026-04124-9
Authors: Simon Liu, Wenjing Zhang, Oded Regev
Institutions: New York University