Back
PubMedPathogenicity prediction

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer.

Nizamuddin N, Biswas S, Zawar A, et al.Front Syst Biol 2026 · August 2026
Relevance score
4/10
Disease / domain
Hereditary cancer
Source
PubMed
PMID 42614771

Tool / method

In silico classification of germline variants by a random forest model built on Ensembl VEP annotations and CADD scores, with no functional work, externally validated on BRCA1 and BRCA2 variants.

Summary

The authors trained a machine learning framework on 104,646 high-confidence ClinVar germline variants (3-star or better review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 scores to separate pathogenic from benign variants. Of four classifiers compared, random forest performed best (AUC-ROC 0.9995; 10-fold cross-validation AUC 0.9992). Applied to 40,894 ClinVar variants of uncertain significance with probability thresholds of 0.80 and 0.20, the model labelled 19,393 (47.4%) as likely pathogenic and 8,957 (21.9%) as likely benign, while 12,544 (30.7%) were conservatively kept uncertain. External validation on 7,462 ENIGMA-classified BRCA1 and BRCA2 variants gave 98.83% concordance, and 89.57% on 671 variants uncertain in ClinVar but resolved by ENIGMA. SHAP-based explainability showed predictions were driven mainly by CADD Phred score, VEP impact tier, consequence class and population allele frequency.

Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.

Analysis

Everything here is in silico and trained on ClinVar labels, with no functional validation at all: none of the relabelled variants is reclassified in clinical practice, where ACMG/AMP rules give computational predictions only limited evidential weight. The high concordance with ENIGMA is partly expected — the easy variants are also the ones CADD separates best — and an AUC of 0.9995 against labels drawn from the same database mainly signals circularity. Any use in clinic would first require calibration against functionally characterised variants and a prospective evaluation of the proposed triage.

Analysis by Dr Thibaut Benquey

Why this score?

Impact 1/3Evidence 1/3Novelty 1/2Sample 1/1Publication 0/1

Clinical impact: 1/3 · Evidence strength: 1/3 · Novelty: 1/2 · Sample size: 1/1 · Publication status: 0/1 → Total: 4/10

Keywords

variant of uncertain significancemachine learningpathogenicity predictionhereditary cancerClinVar
Weekly report in your inbox

Every Wednesday · Annotated selection · Free · Unsubscribe anytime