Publications

You can also find my articles on my Google Scholar profile.

Comprehensive gene heritability estimation reveals the genetic architecture of rare coding variants underlying complex traits

Published in bioRxiv, 2026

Whole-exome sequencing (WES) enables high-resolution interrogation of the contribution of rare coding variants to complex trait variation. However, existing methods for heritability estimation attributed to rare-coding variants are often limited by the effects of linkage disequilibrium (LD) and by the sparse nature of rare variant data. We introduce FLEX (Fast, LD-aware Estimation of eXome-wide and gene-level heritability), a scalable and flexible framework for estimating and partitioning heritability across genes or sets of genes using WES data. FLEX integrates all coding variants– from common to ultra-rare – within a unifled model and corrects for LD-induced effects to improve the accuracy of heritability estimates. In addition, FLEX supports both individual-level and summary statistic data and is computationally efflcient for biobank-scale datasets. FLEX provides an adaptable and accurate approach for quantifying gene-level heritability, advancing our understanding of the genetic architecture of complex traits, and facilitating the discovery of trait-relevant genes.

Authors: Liu, Z., Fu, B., Jeong, M., Anand, A., et al.


Download Paper

Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline

Published in Cell Reports Methods, 2026

Rapid improvements in DNA sequencing technologies have expanded the breadth of genomic features, ranging from nuclear, mitochondrial, and evolutionary variation in germline and somatic contexts, which can be elucidated from sequencing data. In parallel, analytical workflows required to process and identify these features have become increasingly complex, relying on specialized tools and algorithms with varying assumptions and computational requirements. Comprehensive analysis, therefore, requires significant integration effort, limiting scalability, reproducibility, and consistent quality controls. To address this need for a flexible, robust framework that accommodates diverse sequencing methods and feature classes while being highly scalable and adaptable across computational environments, we created metapipeline-DNA to automate genomic analyses.

Authors: Patel, Y., Zhu, C., Yamaguchi, T. N., Wang, N. K., Anand, A., et al.


Download Paper

A biobank-scale method for learning modulators of gene-environment interaction underlying human complex traits from multiple environmental exposures

Published in bioRxiv, 2026

It is increasingly recognized that genetic effects on complex traits and diseases are shaped by environmental context. Biobanks that measure diverse environmental exposures alongside genotypes and phenotypes at scale enable systematic study of gene-environment (G×E) interactions. Existing approaches, however, are limited in their ability to accurately model polygenic G×E involving many exposures across genome-wide genetic variants. It is unclear which exposure combinations are relevant for a given trait while distinguishing true interactions from environment-dependent heteroskedastic noise. To address these challenges, we develop Efficient multi-eNvironmental Gene-environment Interaction iNference Estimator (ENGINE), a supervised variance-component framework that learns an embedding that combines multiple environmental exposures while jointly estimating additive, G×E, and heteroskedastic noise components. To enable biobank-scale inference, ENGINE makes a single pass over the genotype matrix to cache genotype-dependent summaries, then assembles normal-equation components and gradients at each iteration. In simulations, ENGINE controls type I error rates, achieves high power, and accurately recovers the environmental embedding while remaining efficient at biobank-scale. Applied to five complex traits paired with lifestyle exposures in N = 291,273 unrelated white British individuals and M = 454,207 common SNPs (MAF> 0.01) from the UK Biobank, ENGINE recovered G×E variance that was on average 1.4-fold larger than that captured by a single exposure and 5.5-fold larger than that captured by the first principal component of the exposures.

Authors: Liu, Z., Ramteke, A., Anand, A., Gorla, A., Jeong, M., & Sankararaman, S.


Download Paper

A biobank-scale test of marginal epistasis reveals genome-wide signals of polygenic interaction effects

Published in Nature Genetics, 2025

The contribution of genetic interactions (epistasis) to human complex trait variation remains poorly understood due, in part, to the statistical and computational challenges involved in testing for interaction effects. Here we introduce FAME (FAst Marginal Epistasis test), a method that can test for marginal epistasis of a single-nucleotide polymorphism (SNP) on a quantitative trait (whether the effect of an SNP on the trait is modulated by genetic background). FAME is computationally efficient, enabling tests of marginal epistasis on biobank-scale data. Applying FAME to genome-wide association study (GWAS)-significant trait-SNP associations across 53 quantitative traits and ≈300 000 unrelated White British individuals in the UK Biobank (UKBB), we identified 16 significant marginal epistasis signals across 12 traits. Leveraging the scalability of FAME, we further localized marginal epistasis signals across chromosomes and estimated the proportion of variance explained by marginal epistasis effects. Our study provides evidence for interactions between individual genetic variants and polygenic background influencing complex traits.

Authors: Fu, B., Pazokitoroudi, A., Xue, A., Anand, A., Anand, P., Zaitlen, N., & Sankararaman, S.


Download Paper

A scalable adaptive quadratic kernel method for interpretable epistasis analysis in complex traits

Published in Genome Research, 2024

Our knowledge of the contribution of genetic interactions (epistasis) to variation in human complex traits remains limited, partly due to the lack of efficient, powerful, and interpretable algorithms to detect interactions. Recently proposed approaches for set-based association tests show promise in improving power to detect epistasis by examining the aggregated effects of multiple variants. Nevertheless, these methods either do not scale to large numbers of individuals available in Biobank datasets or do not provide interpretable results. We, therefore, propose QuadKAST, a scalable algorithm focused on testing pairwise interaction effects (also termed as quadratic effects) of a set of genetic variants on a trait and quantifying the proportion of phenotypic variance explained by these effects.

Authors: Fu, B.*, Anand, P.*, Anand, A.*, Mefford, J., & Sankararaman, S.


Download Paper