A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer.
Tool / method
In silico classification of germline variants by a random forest model built on Ensembl VEP annotations and CADD scores, with no functional work, externally validated on BRCA1 and BRCA2 variants.
Summary
The authors trained a machine learning framework on 104,646 high-confidence ClinVar germline variants (3-star or better review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 scores to separate pathogenic from benign variants. Of four classifiers compared, random forest performed best (AUC-ROC 0.9995; 10-fold cross-validation AUC 0.9992). Applied to 40,894 ClinVar variants of uncertain significance with probability thresholds of 0.80 and 0.20, the model labelled 19,393 (47.4%) as likely pathogenic and 8,957 (21.9%) as likely benign, while 12,544 (30.7%) were conservatively kept uncertain. External validation on 7,462 ENIGMA-classified BRCA1 and BRCA2 variants gave 98.83% concordance, and 89.57% on 671 variants uncertain in ClinVar but resolved by ENIGMA. SHAP-based explainability showed predictions were driven mainly by CADD Phred score, VEP impact tier, consequence class and population allele frequency.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
Everything here is in silico and trained on ClinVar labels, with no functional validation at all: none of the relabelled variants is reclassified in clinical practice, where ACMG/AMP rules give computational predictions only limited evidential weight. The high concordance with ENIGMA is partly expected — the easy variants are also the ones CADD separates best — and an AUC of 0.9995 against labels drawn from the same database mainly signals circularity. Any use in clinic would first require calibration against functionally characterised variants and a prospective evaluation of the proposed triage.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 1/3 · Evidence strength: 1/3 · Novelty: 1/2 · Sample size: 1/1 · Publication status: 0/1 → Total: 4/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime