dicast: a machine learning method for accurate structural variant detection from short-read sequencing data.
Tool / method
Machine-learning scoring of short-read structural variant calls from alignment and genomic-context features
Summary
Structural variant detection from short-read data remains challenging, even though short-read sequencing underpins most clinical workflows. dicast is a machine-learning method that scores structural variant calls from short-read data using alignment and genomic-context features. The model is trained on a new multi-technology ground truth built from nine samples with extensive manual curation. The authors report that it outperforms existing short-read callers and consensus approaches, recovering substantially more true positives at high precision. Applied to diagnostics, dicast identified all pathogenic variants across multiple disease cohorts and 20% more candidate pathogenic deletions than consensus approaches.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
What matters for a laboratory is that this is not about moving to long-read but about getting more out of short-read data already generated — structural variants remain the worst-served class in exome and genome reporting. The ground truth rests on nine samples, however, and the extra 20% of deletions are candidates, not diagnoses: that is additional interpretation work, and the abstract does not say what fraction is confirmed. Worth benchmarking on in-house data before any pipeline change.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 3/3 · Evidence strength: 2/3 · Novelty: 2/2 · Sample size: 0/1 · Publication status: 1/1 → Total: 8/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime