1,721,227 research outputs found
Recommended from our members
Model-driven optimization of high-throughput in vivo CRISPR screen design
The recent developments of the CRISPR/Cas9 gene-editing system have made way for large-scale, loss-of-function genetic screens that can identify genes underlying a given phenotype, known as high-throughput CRISPR screens. By leveraging the precision of CRISPR/Cas9 and the capacity to capture millions of cells in one library preparation, these screens enrich and deplete the expression of various specific genes, identifying gene functions that help elucidate genotype-phenotype relationships. Furthermore, by modulating genetic interactions, these screens can uncover gene regulatory mechanisms, revealing genetic dependencies. Although these screens have shown to be incredibly effective, they are often prohibitively expensive. Additionally, there is a lack of information and tools to determine the optimal experimental design, such that the most informative data is produced, given experimental constraints.Here, we introduce a statistical model that simulates high-throughput in vivo CRISPR screens to provide insight into optimizing the experimental protocol. We first demonstrate our model successfully simulates such screens by comparing the generated data with real experimental data. Then, we simulate screens across varying parameter inputs and investigate their influence on statistical power. Given our findings, we conclude with general guidelines and suggestions for effectively designing high-throughput in vivo CRISPR screens
Recommended from our members
Methods for the Quantitative Characterization of the Genetic Basis of Human Complex Traits
A major finding from the last decade of genome-wide association studies (GWAS) is that variant-phenotype associations are significantly enriched in noncoding regulatory regions of the genome. This result suggests that GWAS associations localize variants that modulate phenotype via gene regulation as opposed to alterations in protein structure/function. However, for most complex traits, most aspects of genetic architecture—the number of causal variants/genes for a trait and the degree to which causal effect sizes are coupled with genomic features such as minor allele frequency (MAF) and linkage disequilibrium (LD)—remain actively debated. In this dissertation, I introduce three new methods to explore and quantitatively characterize complex-trait genetic architecture. First, I derive an unbiased estimator of genome-wide SNP-heritability under a very general random effects model that makes minimal assumptions on the underlying (unknown) genetic architecture of the trait. Second, I introduce a method for estimating the number of causal variants that are shared between two ancestral populations for a given trait, and I discuss the implications of the method and real-data results for improving polygenic risk prediction in ethnic minority populations. Third, I propose methods for partitioning the heritability of individual genes by MAF to identify disease-relevant genes, with the hypothesis that some disease-relevant genes may have relatively large heritability contributions from rare and low-frequency variants while still having low total gene-level heritability
Recommended from our members
Computational methods to analyze large-scale genetic studies of complex human traits
Large-scale genome-wide association studies (GWAS) have produced a rich resource of genetic data over the past decade, urging the need to develop computational and statistical methods that analyze these data. This dissertation presents four statistical methods that model the correlation structure between genetic variants and its effect on GWAS summary association statistics to help understand the genetic basis of complex human traits and diseases.The first method employs the multivariate Bernoulli distribution to model haplotype data, allowing for higher-order interactions among genetic variants, and shows better accuracy in predicting DNase I hypersensitivity status.The second method partitions heritability into small regions on the genome using GWAS summary statistics data, while accounting for complex correlation structures among genetic variants, and uncovers the genetic architectures of complex human traits and diseases.Extending the second method into pairs of traits, the third method partitions genetic correlation into small genomic regions using GWAS summary statistics data, and provides insights into the shared genetic basis between pairs of traits.Finally, the fourth method dissects population-specific and shared causal genetic variants of complex traits in two continental populations, using GWAS summary statistics data obtained from samples of different ethnicities, and reveals differences in genetic architectures of two continental populations
Recommended from our members
Uncertainty, portability and ancestry in polygenic scoring
Polygenic score (PGS) is a tool for understanding an individual's predisposition to certain diseases or complex traits based on its genetic profile. In the burgeoning era of genomic medicine, PGS has emerged as a promising tool in advancing precision healthcare, demonstrating versatile utility such as patient risk stratification, disease risk prediction, and disease subtyping. However, its real application in clinical settings is limited by its uncertainty, bias, and low portability across diverse populations. For example, an individual may receive different genetic risk reports from different providers, and the score for a non-European individual may be less accurate than for a European individual. To fully understand and partially address these limitations, I first developed a Bayesian method to quantify the uncertainty in PGS at the individual level. I find trait-specific genetic architecture such as larger polygenicity and lower heritability combined with a small training sample size will lead to large uncertainty in PGS estimate, which in turn results in unreliable patient stratification in downstream analysis. Next, I expanded this approach to encompass individuals from varied genetic ancestry backgrounds. I find that the PGS performance varied from individual to individual with genetic distance playing a key role in impacting the performance of PGS; larger genetic distance from training data correlates with higher uncertainty and lower accuracy in testing individuals. These findings highlight the necessity of integrating individual-level PGS metrics in personalized medicine and the need for increasing genetic research diversity to ensure equitable and responsible use of PGS in clinical settings
Low-coverage transcriptomics for understanding genetic regulation of complex traits
Mapping genetic variants that regulate gene expression (eQTL mapping) in large-scale RNA sequencing (RNA-seq) studies is often employed to understand functional consequences of regulatory variants. However, the high cost of RNA-Seq limits sample size, sequencing depth, and therefore, discovery power in eQTL studies. In this work, we demonstrate that, given a fixed budget, eQTL discovery power can be increased by lowering the sequencing depth per sample and increasing the number of individuals sequenced in the assay. We perform RNA-Seq of whole blood tissue across 1490 individuals at low-coverage (5.9 million reads/sample) and show that the effective power is higher than that of an RNA-Seq study of 570 individuals at moderate- coverage (13.9 million reads/sample). Next, we leverage synthetic datasets derived from real RNA-Seq data (50 million reads/sample) to explore the interplay of coverage and number individuals in eQTL studies, and show that a 10-fold reduction in coverage leads to only a 2.5- fold reduction in statistical power to identify eQTLs. Our work suggests that lowering coverage while increasing the number of individuals in RNA-Seq is an effective approach to increase discovery power in eQTL studies. We then build a pipeline using existing tools CIBERSORTx and bMIND to computationally deconvolute low-coverage bulk RNA-seq from a total of 1,996 individuals to estimate cell type expression. We show that cell type expression estimates are consistent with those from scRNA-seq and can be used as a powerful approach to finding ct- eQTLs. Next, we use medication history from this cohort to look for SNP x lithium interactions in ct-eQTLs, finding 110 examples of eGenes whose cell type expression is significantly associated with some SNP dependent of lithium usage
Recommended from our members
Genetic mapping, inference and prediction across diverse human populations
Genome-wide association studies have revolutionized our understanding of genetic influences on common diseases and complex traits. However, the majority of discoveries have been limited to individuals of European ancestry, leading to a data collection bias that disproportionately under-samples non-European populations. This bias leads to missed discovery opportunities and differential prediction accuracy across sub-populations defined by genetic ancestry and socioeconomic factors. Although datasets with diverse genetic ancestry backgrounds are increasingly available, existing analytical tools often fail to account for the heterogeneity present in these datasets. Here, I introduce new computational and statistical methods for genetic mapping, inference, and prediction across diverse human populations. First, I investigate the power of genetic mapping approaches in populations with diverse genetic ancestry backgrounds. Second, I explore the inference of genetic architecture, estimating the cross-ancestry sharing of genetic effects. Third, I examine genetic prediction, quantifying differential polygenic scoring accuracy by contexts and developing an approach to account for such differences
Recommended from our members
Integrative statistical methods to understand the genetic basis of complex trait
The Genome-wide Association study (GWAS) is one of the primary tools for understanding the genetic basis of complex traits. In this dissertation I introduce enhanced statistical methods to do integrative GWAS analysis with functional genomic data. First, I describe an integrative fine-mapping framework to prioritize causal variants at known GWAS risk loci. Next, I expand upon this framework to exploit genetic heterogeniety across human populations to improve statistical efficiency. I then consider a new inference strategy to reduce the computational burden of the methodology. Finally, I propose a new approach for GWAS discovery that leverages functional genomic data through polygenic modeling
Recommended from our members
Exploring the roles of genetic regulation in human phenotypes
Human phenotypes are influenced to varying extents by inherited genetic variation, although specific mechanisms through which this variation affects the phenotypes are not completely understood. In this dissertation I explore different modes of genetic regulation in the context of human complex traits and rare disorders. First, I examine the degree of shared genetic basis between complex traits and rare monogenic disorders across a wide range of phenotypes. Second, I explore the regulatory landscape of ovarian surface epithelial cells to identify putative pathways involved in the development of epithelial ovarian cancer. This work provides a foray into understanding the different ways that genetic variation can drive downstream phenotypes through direct and epigenetic regulation of target genes
Recommended from our members
Computational methods to discern the genetic basis of complex disease
Genome-wide association studies (GWAS) have identified thousands of regions in the genome containing risk variants for complex traits. Due to the correlation structure between ge- netic variants, there is a need for computational methods that can tease apart causal from non-causal variants in these implicated regions. This dissertation presents three statistical methods that aim to improve our detection of causal variants at risk regions and ultimately better our understanding of the genetic basis of complex disease.The first method aims to fine-map genetic regions impacting multiple correlated traits at once, employing the Multivariate Normal (MVN) distribution to jointly model association statistics at a risk region.The second method performs hierarchical fine-mapping on risk regions that show evidence for a SNP impacting gene expression through an epigenetic feature, such as histone modifi- cations. It uses both the MVN as well as the Matrix-variate Normal distribution to jointly model effects from SNP to epigenetic mark to gene expression.The third method builds on existing summary statistics imputation methods by integrating functional annotation data to improve prediction of associations at untyped SNPs
Recommended from our members
Methods for Optimizing Mechanistic and Predictive Models of Human Disease
A major goal of the biomathematics discipline is to optimize mathematical models for biological processes. This optimization can take on various forms; finding the appropriate model that fits available data, allows for accurate inference, and is computationally feasible is no easy task and requires an understanding of both the biological processes at hand and the mathematics behind each potential model or algorithm. In this dissertation, I seek to understand how mathematical modeling choices affect our ability to understand human disease. I study infectious, cancerous, and polygenic disease from a variety of computational perspectives. First, I apply methods for differential sensitivity analysis in biological models for both cancerous and infectious disease spread. I compare prediction accuracy for existing first-order methods and propose a second-order method with enhanced flexibility both in terms of the model for which it is applied and the programming environment available. Second, I compare statistical approaches for uncovering genetics of complex disease in admixed populations, using likelihood ratio tests to understand how to incorporate local ancestry in genome wide association studies to achieve the highest power. Third, I utilize machine learning methods to reduce diagnostic delay for patients across the University of California Health system. I adapt a logistic regression model to find patients likely to have common variable immune deficiencies from one health system to five health systems. I also adapt this algorithm from the immunology realm to the cardiology realm to predict cardiac amyloidosis. Along the way, I use this context to study automated feature selection, longitudinal feature engineering, and observational bias in electronic health record data
- …
