Conferences

A scalable adaptive quadratic kernel method for interpretable epistasis analysis in complex traits

Presented at Research in Computational Molecular Biology (RECOMB Main Conference), 2024

Our knowledge of the contribution of genetic interactions (epistasis) to variation in human complex traits remains limited, partly due to the lack of efficient, powerful, and interpretable algorithms to detect interactions. Recently proposed approaches for set-based association tests show promise in improving power to detect epistasis by examining the aggregated effects of multiple variants. Nevertheless, these methods either do not scale to large numbers of individuals available in Biobank datasets or do not provide interpretable results. We, therefore, propose QuadKAST, a scalable algorithm focused on testing pairwise interaction effects (also termed as quadratic effects) of a set of genetic variants on a trait and quantifying the proportion of phenotypic variance explained by these effects.

Authors: Fu, B.*, Anand, P.*, Anand, A.*, Mefford, J., & Sankararaman, S.


Download Paper

Explaining potential epistasis in genomic data using symbolic representations of complex black box models

Presented at Bruins in Genomics, 2022

Epistasis, known as the interaction among genetic variants, has long been hypothesized to play a major role in explaining missing heritability. Though recent studies have found many candidate variants demonstrating epistasis signals in the UKBiobank, it remains a controversial question how to interpret the findings. Nonlinear models have shown potential in capturing these signals, but we require additional explanation methods to understand the projected relationships. Here, we utilize symbolic pursuit, a form of symbolic regression that provides a closed-form, interpretable model which generalizes first order explanations. Furthermore, we extend this study by applying Taylor expansions to the model, balancing interpretability with performance while improving its generalizability. We found the method performed reliably and was consistent with other methods across a variety of simulated data. This work contains strong implications for its use on large genomic datasets and its ability to capture nonlinear interactions without prior knowledge of the genetic architecture.

Authors: Anand, A., Anand, P.


Download Paper