Interpretable distillation reveals that deep learning splicing models suffer from pervasive confounders and blind spots
Abstract
BACKGROUND: Predicting RNA splicing from genomic sequence is a crucial task for understanding gene regulation and interpreting genetic variation. Recent deep learning advancements have led to splicing prediction algorithms that achieve state-of-the-art performance compared to earlier models. However, owing to the limited interpretability of deep learning models, the predictive mechanisms of current splicing models remain poorly understood. RESULTS: Here we develop a framework to explain model prediction logic using interpretable distillation. Applying our framework, we find that RNA splicing prediction models suffer from pervasive confounders and blind spots, leading to poor performance on non-reference sequences. We find that splicing models recognize exons through surprisingly simple additive combinations of sequence motifs, including known splicing regulatory elements. Critically, our analysis also reveals that splicing models exploit genomic confounders unrelated to splicing and fail to adequately capture the effects of RNA structure, leading to systematic prediction errors. CONCLUSIONS: Our findings illuminate fundamental limitations of training models on genomic sequences and suggest ways to overcome them.
// Source
Authors: Simon Liu, Wenjing Zhang, Oded Regev
Institutions: New York University