1,721,129 research outputs found

    A Transcriptomic Atlas of the Human Brain Reveals Genetically Determined Aspects of Neuropsychiatric Health

    No full text
    AbstractImaging features associated with neuropsychiatric traits can provide valuable insights into underlying pathophysiology. Using data from the UK biobank, we perform tissue-specific TWAS on over 3,500 neuroimaging phenotypes to generate a publicly accessible resource detailing the neurophysiologic consequences of gene expression. As a comprehensive catalog of neuroendophenotypes, this resource represents a powerful neurologic gene prioritization schema that can improve our understanding of brain function, development, and disease. We show that our approach leads to highly reproducible results. Notably, genetically determined expression alone is shown here to enable high-fidelity reconstruction of brain structure and organization. We demonstrate complementary benefits of cross-tissue and single-tissue analyses towards an integrated neurobiology and provide evidence that gene expression outside the central nervous system provides unique insights into brain health. As an application, we show that over 40% of genes previously associated with schizophrenia in the largest GWAS meta-analysis causally affect neuroimaging phenotypes noted to be altered in schizophrenic patients.</jats:p

    Multilayer modelling of the human transcriptome and biological mechanisms of complex diseases and traits

    Get PDF
    Abstract: Here, we performed a comprehensive intra-tissue and inter-tissue multilayer network analysis of the human transcriptome. We generated an atlas of communities in gene co-expression networks in 49 tissues (GTEx v8), evaluated their tissue specificity, and investigated their methodological implications. UMAP embeddings of gene expression from the communities (representing nearly 18% of all genes) robustly identified biologically-meaningful clusters. Notably, new gene expression data can be embedded into our algorithmically derived models to accelerate discoveries in high-dimensional molecular datasets and downstream diagnostic or prognostic applications. We demonstrate the generalisability of our approach through systematic testing in external genomic and transcriptomic datasets. Methodologically, prioritisation of the communities in a transcriptome-wide association study of the biomarker C-reactive protein (CRP) in 361,194 individuals in the UK Biobank identified genetically-determined expression changes associated with CRP and led to considerably improved performance. Furthermore, a deep learning framework applied to the communities in nearly 11,000 tumors profiled by The Cancer Genome Atlas across 33 different cancer types learned biologically-meaningful latent spaces, representing metastasis (p < 2.2 × 10−16) and stemness (p < 2.2 × 10−16). Our study provides a rich genomic resource to catalyse research into inter-tissue regulatory mechanisms, and their downstream consequences on human disease

    Modelling of the transcriptome using networks

    No full text
    This repository contains all the code necessary to run and further extend the experiments presented in our paper. Please cite: Azevedo, Tiago, Dimitri, Giovanna Maria, Lio, Pietro, & Gamazon, Eric R. (2020) "Multilayer modelling and analysis of the human transcriptome." bioRxiv. https://www.biorxiv.org/content/10.1101/2020.05.21.109082v4 Azevedo, Tiago, Dimitri, Giovanna Maria, Lio, Pietro, & Gamazon, Eric R. (2020, May 25). Modelling of the transcriptome using networks (Version 1). Zenodo. http://doi.org/10.5281/zenodo.3842659 Abstract In the present work, we performed a comprehensive intra-tissue and inter-tissue network analysis of the human transcriptome. We generated an atlas of communities in co-expression networks in each of 49 tissues and evaluated their tissue specificity. UMAP embeddings of gene expression from the identified communities recovered biologically meaningful tissue clusters, based on tissue organ membership or known shared function. We developed an approach to quantify the conservation of global structure and estimate the sampling distribution of the distance between tissue clusters via bootstrapped manifolds. We found not only preserved local structure among clearly related tissues (e.g., the 13 brain regions) but also a strong correlation between the clustering of these related tissues relative to the remaining ones. Interestingly, brain tissues showed significantly higher variability in community size than non-brain (p = 1.55x10-4). We identified communities that capture some of our current knowledge about biological processes, but most are likely to encode novel and previously inaccessible functional information. For example, we found a 17-member community present across all of the brain regions, which shows significant enrichment for the nonsense-mediated decay pathway (adjusted p = 1.01x10-37). We also constructed multiplex architectures to gain insights into tissue-to-tissue mechanisms for regulation of communities in the transcriptome, including communities that are likely to play a functional role throughout the central nervous system (CNS) and communities that may participate in the interaction between the CNS and the enteric nervous system. Notably, new gene expression data can be embedded into our models to accelerate discoveries in high-dimensional molecular datasets. Our study provides a rich resource of co-expression networks, communities, multiplex architectures, and enriched pathways in a broad collection of tissues, to catalyse research into inter-tissue regulatory mechanisms and enable insights into their downstream phenotypic consequences

    COMBINING METHODOLOGIES FOR IMPROVING INTERPRETABILITY AND GENERALIZABILITY OF GENOMIC INVESTIGATIONS

    No full text
    Paralleling the expansive growth of large-scale biobanks that contain patient electronic medical records-tied genome wide genetic data, the number of genome-wide association studies (GWASs) and the volume of newly identified disease associated genetic factors have expectedly accelerated in size as well. While this has facilitated an ongoing discovery process of countless disease-associated genetic loci, there remains an enormous gap in understanding the biological consequences and mechanisms related to these genomic areas. To address this gap, many bioinformatic methods including colocalization with functional annotations have been developed to better interpret GWAS findings. In this manner, many methods help identify expression quantitative trait loci (eQTLs), areas of the genome shown to regulate gene expression. Advancing this endeavor even further, several efforts have trained imputed gene expression models that capture the cumulative cis-regulatory effects on gene expression. Because of the immediate interpretability of these models, they can be utilized in regression models for transcriptomic wide association studies (TWAS) to characterize specific gene associations with phenotypes and diseases of interest. The work presented here largely strives to leverage this feature of interpretability by integrating these imputed gene expression models with other modeling methods to help design more interpretable genomic studies. By integrating with Bayesian frameworks, linear mixed effects models, and additive models, both unbiased and hypothesis-driven methods are demonstrated and characterized in the context of both common and rare diseases. This versatile collection of tools hopes to provide diverse means of approaching clinical phenotype by allowing for interrogation of both the entire genome and specific functional partitions of genetic factors. In addition to these transcriptomic-based tools, a method of quantifying cumulative ancestry-driven genetic effects on phenotypes is also demonstrated to help directly in interpretation of existing GWAS studies. Taken together, these methodologies both create approaches to better interpret existing results as well as provide tools for designing genomic investigations that possess immediate interpretability

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Integrating Large-Scale Human Genetic and Regulatory Genomic Data to Functionally Annotate ctcf Binding Variation

    No full text
    CCCTC binding factor (CTCF) regulates gene expression through DNA binding at thousands of genomic loci. Genetic variation in these CTCF binding sites (CBSs) are important drivers of phenotypic variation, yet extracting those that are likely to have functional consequences in whole genome sequencing (WGS) remains challenging. Through this dissertation, I explore conceptual frameworks to identify and prioritize CBS variants in gnomAD, a WGS database consisting of 76,156 individuals. First, I integrate computational and experimental predictions of CTCF binding into an empirical false-positive measure that can be applied to the score distribution of a precision-weight matrix. I then synthesize CTCF’s binding patterns at 1,063,878 genomic loci across 214 biological contexts into a summary of binding activity. This measure correlates with both conserved nucleotides and sequences that contain high-quality CTCF binding motifs. Finally, I use binding activity to evaluate high confidence allelic binding predictions for 1,253,329 SNVs in gnomAD that disrupt a CBS. I find a strong, positive relationship between the mutability adjusted proportion of singletons (MAPS) metric and the loss of CTCF binding at loci with high in vitro activity. Together, this body of work nominates thousands of rare, noncoding variants that disrupt CTCF binding for further functional studies while providing a blueprint for synthesizing large-scale genomic data to better prioritize noncoding variation in human disease studies
    corecore