Research program

How do we determine the significance of a genetic variant?

Integrating in vitro and in silico evidence to quantify disease probability rather than relying on binary pathogenicity labels.

What is a variant’s significance to those who carry it?

Sequencing reveals far more variants than anyone can interpret one at a time, and most rare alleles end up “of uncertain significance.” We resolve them by integrating published carriers, functional measurements, protein structure and computational predictors into a probability of disease with explicit uncertainty. Cardiac ion-channel genes for long QT, Brugada and CPVT are our starting point because they connect molecular mechanism to a clinical outcome cleanly. For an overview, watch this talk.

1,400+

SCN5A variants curated

5

evidence sources unified

LLM consensus per curation

4

genes live on VariantBrowser

How the pipeline works

An open-source path from the published record to a penetrance estimate.

01

Curate

LLM agents pull carrier counts, phenotypes and study context from PubMed; a multi-model consensus scores gene–disease validity against the ClinGen SOP. Every extraction links back to its sentence in the source.

02

Annotate

AlphaMissense, REVEL, CADD, ClinVar and gnomAD are unified into one database so every variant carries the same panel of predictive evidence.

03

Localize

Variants are mapped onto AlphaFold structures with elastic-network analysis, so pathogenic clusters raise the prior for nearby variants of uncertain significance.

ProteinProximityAnalysis · in development
04

Estimate

A Bayesian model fuses carriers, features and structural priors into a penetrance estimate with full uncertainty. Patented as US-20220406461-A1.

Variant-level vs. individual

Penetrance is a property of the variant. Risk is a property of the patient.

Variant-level estimate. The probability that a heterozygote carrying this variant is affected, averaged over all carriers. This is what VariantBrowser reports, and it is calibrated against published cohorts.

Individual clinical risk. The same variant produces different outcomes in different people because of polygenic background, sex, QTc, age and therapy. Estimating that requires patient-derived models and longitudinal cohorts, which is where the newer projects below sit.

Perturbation & disease

How much functional perturbation is too much?

Across more than 1,400 curated SCN5A variants, including 304 with functional data, modest perturbations produce heterogeneous presentations while extreme loss or gain of function produces consistent ones. Carriers of the same variant can still present differently; peak current tracks with, but does not determine, how many are diagnosed with Brugada syndrome.

VariantPeak currentUnaffectedBrS1
S1787N95%121
Y1795H66%75
R367H0%316

Peak current is a proxy for channel function; counts are published heterozygotes.

Disease-associated variants in Nav1.5 mapped onto channel structure
Per-residue rates of Brugada syndrome (BrS1), type-3 long QT (LQT3) and unaffected carriers cluster differently across NaV1.5 membrane regions.
High-throughput characterization

Measuring variant effects at the scale sequencing demands.

Deep mutational scanning

The most complete variant-effect map to date for KCNH2/Kv11.1 trafficking, thousands of variants in one experiment, and a contribution to the CardioVar atlas of variant effects across cardiovascular genes (LDLR, KCNQ1).

Calibrated automated patch clamp

With the Vandenberg and Ng labs, a calibrated PS3/BS3 assay for KCNH2 that measurably lowers clinical uncertainty and, combined with MAVE data, improves cardiac-event risk stratification (Circulation 2024).

iPSC-cardiomyocytes

Patient-derived lines from individuals at the extremes of QT polygenic score, CRISPR-edited rare variants, and an open field-potential analyzer, used to test how genetic background reshapes channel function and drug response.

All-atom molecular dynamics of a channel segment in a lipid bilayer
All-atom molecular dynamics (AMBER) of a channel segment in an explicit lipid bilayer.
Structure & mechanism

From polygenic background to molecular mechanism.

Quantitative proteomics in iPSC-cardiomyocytes shows how elevated QT polygenic risk reshapes the Kv11.1 (hERG) interaction network around endocytic trafficking, a concrete mechanism linking common variation to repolarization. Rosetta modeling, molecular dynamics and NMR supply the structural priors the penetrance model uses.

Where the program is going

Patient-centered outcome prediction

With Bastarache and Ruderfer (NHGRI UG3), integrating multimodal data with ML/AI to predict the outcomes that matter most to people who receive a pathogenic result.

Breakthrough events on therapy

We are evaluating survival models in harmonized international KCNH2 cohorts to test whether variant- and patient-specific features can identify carriers at higher risk of cardiac events despite beta-blockers.

Beyond channelopathies

The framework is gene-agnostic. Pilots extend LLM-assisted, ClinGen-style curation to atrial fibrillation genes and toward 50 gene–disease pairs with structure-derived features.

Funding

NIH/NHLBI R01HL160863

Integrating KCNH2 variant-specific features and heterozygote phenotypes to estimate long QT penetrance. PI, 2022–2027.

NIH/NHLBI R01HL164675 · CardioVar

Systematically mapping variant effects for more than 25 cardiovascular disease genes. Key personnel (Roden, PI), 2022–2026.

NIH/NHGRI UG3HG014376

Patient-centered prediction of clinically important outcomes arising from pathogenic variants. Co-I (Bastarache/Ruderfer, PIs), 2025–2027.

Stanley Cohen Award for Genomic Discovery

From QT polygenic extremes to biological mechanisms: a BioVU functional genomics pilot. PI, VUMC, 2026–2027.

Previously: NIH K99/R00 HL135442 (2017–2022), AHA Career Development Award (2021–2024), Leducq Foundation 18CVD05 (2019–2024), NIH R01HL149826 (2020–2023).

National Heart, Lung, and Blood Institute U.S. Department of Health and Human Services Vanderbilt University Medical Center