Genomic language model for predicting enhancers and their allele-specific activity in the human genome.
Tool / method
DNABERT language model fine-tuned on ENCODE candidate cis-regulatory elements, combined with transcription factor models to predict allele-specific effects
Summary
DNABERT-Enhancer applies the pre-trained DNABERT language model to enhancer prediction in the human genome. The benchmark dataset was built from the ENCODE registry of candidate cis-regulatory elements: 21,926 enhancers of 201 bp and 46,159 enhancers of 350 bp as positive instances. The best fine-tuned model achieved 88.05% accuracy and a Matthews correlation coefficient of 76% on an independent dataset; applied genome-wide it identified 1,684,595 enhancer regions covering 26.65% of the human genome. Integrative analysis with DNABERT transcription factor models identified 2,681 loss-of-function and 1,917 gain-of-function enhancer variants, altering 1,623 and 1,247 ENCODE cCRE enhancers respectively, plus 4,057 candidate de novo enhancers created by 5,464 gain-of-function variants. Code, trained models and an interactive web application are freely available.
Synthesis written by Geno'X. For the full original abstract, please refer to the source publication.
Analysis
Non-coding sequence remains the blind spot of exome and above all genome interpretation, and a variant-level enhancer catalogue is useful in itself. But the evaluation is done entirely on ENCODE data, with no patients: nothing establishes that a variant labelled loss-of-function here has a phenotypic effect, and 26.65% of the genome annotated as enhancer is far too broad to filter a diagnostic genome. Use it as an annotation resource, not as classification evidence.
Analysis by Dr Thibaut Benquey
Why this score?
Clinical impact: 1/3 · Evidence strength: 1/3 · Novelty: 2/2 · Sample size: 1/1 · Publication status: 1/1 → Total: 6/10
Keywords
Every Wednesday · Annotated selection · Free · Unsubscribe anytime