ME AI LGAug 18, 2025

A Generalized Genetic Random Field Method for the Genetic Association Analysis of Sequencing Data

Ming Li, Zihuai He, Min Zhang, Xiaowei Zhan, Changshuai Wei, Robert C Elston, Qing Lu

arXiv:2508.12617v11.213 citationsh-index: 88

Originality Incremental advance

AI Analysis

This work addresses the need for advanced statistical methods to analyze sequencing data in genetics, particularly for rare variants, but it is incremental as it builds on existing similarity-based approaches.

The authors tackled the challenge of analyzing high-dimensional sequencing data for genetic association studies by proposing a generalized genetic random field (GGRF) method, which demonstrated improved or comparable power to SKAT in simulations, especially for rare variants, and successfully detected associations of two genes with serum triglyceride in a real dataset.

With the advance of high-throughput sequencing technologies, it has become feasible to investigate the influence of the entire spectrum of sequencing variations on complex human diseases. Although association studies utilizing the new sequencing technologies hold great promise to unravel novel genetic variants, especially rare genetic variants that contribute to human diseases, the statistical analysis of high-dimensional sequencing data remains a challenge. Advanced analytical methods are in great need to facilitate high-dimensional sequencing data analyses. In this article, we propose a generalized genetic random field (GGRF) method for association analyses of sequencing data. Like other similarity-based methods (e.g., SIMreg and SKAT), the new method has the advantages of avoiding the need to specify thresholds for rare variants and allowing for testing multiple variants acting in different directions and magnitude of effects. The method is built on the generalized estimating equation framework and thus accommodates a variety of disease phenotypes (e.g., quantitative and binary phenotypes). Moreover, it has a nice asymptotic property, and can be applied to small-scale sequencing data without need for small-sample adjustment. Through simulations, we demonstrate that the proposed GGRF attains an improved or comparable power over a commonly used method, SKAT, under various disease scenarios, especially when rare variants play a significant role in disease etiology. We further illustrate GGRF with an application to a real dataset from the Dallas Heart Study. By using GGRF, we were able to detect the association of two candidate genes, ANGPTL3 and ANGPTL4, with serum triglyceride.

View on arXiv PDF

Similar