Interpretable distillation reveals that deep learning splicing models suffer from pervasive confounders and blind spots.
Tool / method
Exon recognition through additive combinations of sequence motifs, with reliance on genomic confounders unrelated to splicing and inadequate handling of RNA structure.
Summary
Predicting splicing from genomic sequence is central to interpreting genetic variation, and recent deep learning models achieve state-of-the-art performance, yet their prediction logic remains poorly understood. The authors develop a framework that explains model prediction logic through interpretable distillation and apply it to existing splicing models. They show that these models recognise exons through surprisingly simple additive combinations of sequence motifs, including known splicing regulatory elements. Critically, the analysis reveals that the models exploit genomic confounders unrelated to splicing and fail to capture the effects of RNA structure, producing systematic prediction errors and poor performance on non-reference sequences. The authors conclude that training on genomic sequences carries fundamental limitations and suggest ways to overcome them.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
This bears directly on variant interpretation: a predicted splicing score weighs heavily when classifying a candidate variant, and the paper shows that these models degrade precisely on non-reference sequences, that is, on the very sequence carrying the variant under evaluation. The sensible conclusion is not to discard the predictors but to treat an isolated score as weak evidence until it is backed by functional proof from patient RNA or a minigene assay, especially for noncanonical variants. What is missing is a quantification of the error on a real series of clinical variants: as it stands we know the flaw exists but not in which situations the score remains usable.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 2/3 · Evidence strength: 2/3 · Novelty: 2/2 · Sample size: 0/1 · Publication status: 1/1 → Total: 7/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime