Early identification of rare disease using deep phenotyping of electronic health records
Tool / method
Machine-learning models trained on billing codes and on phenotypes extracted from clinical notes
Summary
Patients with rare genetic disorders often wait years for a correct diagnosis. The authors trained machine-learning models on billing codes and on phenotypes extracted from clinical notes, then applied them to eight conditions in a longitudinal cohort of roughly 3 million patient records from the Mayo Clinic. The proportion of patients identified early varied widely by condition, from 10% to 89% at 99% specificity. Median lead times over the first diagnostic billing code exceeded one year for most conditions: 1,408 days for Fabry disease, 2,755 days for hereditary angioedema, 1,067 days for neurofibromatosis type 1 and 2,320 days for hereditary hemorrhagic telangiectasia. Integrating diverse phenotypic data sources provided superior predictive value.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
Case-finding from health records targets the right bottleneck: diagnostic odyssey usually stems from not thinking of the test rather than from the test's performance. Two major caveats: the 10% to 89% spread across conditions rules out reasoning from a global figure, and 99% specificity applied to a rare disease still generates a large number of false positives to review. A single-centre preprint on retrospective data whose label is a billing code: prospective validation remains entirely to be done.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 3/3 · Evidence strength: 2/3 · Novelty: 1/2 · Sample size: 1/1 · Publication status: 0/1 → Total: 7/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime