A Reproducible MFASS Benchmark of Splice-Disruption Predictors Reveals a Shared Exon-Interior Blind Spot
Tool / method
Reproducible benchmark of five splicing predictors against a multiplexed splicing assay (MFASS), with performance stratified by distance to the splice site.
Summary
The authors benchmark four published splicing variant-effect predictors against a multiplexed experimental splicing assay, on 27,733 single-nucleotide variants in and around human exons from MFASS with measured exon-inclusion outcomes. Pangolin is the strongest predictor of splice-disrupting variants (AUROC 0.888, average precision 0.421), ahead of SpliceAI (0.819; 0.321) and SpliceTransformer (0.786; 0.317), with MMSplice fourth (0.758; 0.256); all four exceed the older SPANR model (0.748; 0.228). A calibrated consensus of the three deep-learning sequence-window predictors, evaluated on an exon-grouped held-out split, does not meaningfully improve over Pangolin alone. Stratifying by distance to the splice site exposes a shared blind spot: all tools detect disruptions within a few bases of the splice site well, but recall declines sharply in the exon interior, and 19% of disrupting variants are missed by every tool, these shared misses being predominantly exon-interior. MMSplice, the one model built for modular exonic and intronic effects rather than splice-site recognition, shows the same distance-dependent decline, so the blind spot is not an artefact of splice-site-centric architectures.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
The blind spot documented here — 19% of disrupting variants missed by all five tools, predominantly inside exons — sits precisely on the ground claimed by SPiP (Splicing Prediction Pipeline; Leman R, et al. Hum Mutat 2022;43(12):2308-2323), a French tool absent from the benchmark and built in response to the fact that most in silico predictors target a single type of splicing motif, whereas SPiP assesses 5' and 3' sites, branch points and splicing regulatory elements in a single machine learning analysis. The caveat is essential: MMSplice, the only model in the benchmark designed for modular exonic and intronic effects, shows exactly the same decline with distance to the splice site — being built for exonic effects therefore does not guarantee closing the blind spot, and the question remains open for SPiP. The figures above all are not comparable: the published AUC of 0.986 for SPiP (83.13% sensitivity, 99% specificity, versus 0.965 for SpliceAI and 0.766 for SQUIRLS) comes from its own set of 4,616 curated variants across 227 genes with RNA studies, whereas the AUROC values in this benchmark come from a multiplexed assay on 27,733 variants; placing 0.986 next to 0.888 would be a reasoning error. The merit of this work is precisely to make the question testable: the code and dataset are public and reproducible, so any team can add SPiP and settle it.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 2/3 · Evidence strength: 3/3 · Novelty: 1/2 · Sample size: 1/1 · Publication status: 0/1 → Total: 7/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime