1,720,997 research outputs found

    Designing a Robust and Portable Workflow for Detecting Genetic Variants Associated with Molecular Phenotypes Across Multiple Studies

    Get PDF
    Kvantitatiivse tunnuse lookusteks (quantitative trait locus, QTL) nimetatakse geneetilisi variante, millel on statistiline seos mõne molekulaarse tunnusega. QTL analüüs võimaldab paremini aru saada komplekshaiguseid ja tunnuseid mõjutavatest molekulaarsetest mehhanismidest. Tüüpiline QTL analüüs koosneb suurest hulgast sammudest, mille kõigi jaoks on olemas palju erinevaid tööriistu, kuid mida ei ole siiani kokku pandud ühte lihtsasti kasutatavasse, teisaldatavasse ning korratavasse töövoogu. Käesolevas töös loodud töövoog koosneb kolmest moodulist: huvipakkuva tunnuse kvantifitseerimine (i), andmete normaliseerimine ja kvaliteedikontroll (ii) ning QTL analüüs (iii). Kvantifitseerimise ja QTL analüüsi moodulite jaoks kasutasime Nextflow töövoo juhtimise süsteemi ning järgisime kõiki nf-core raamistiku parimaid praktikaid. Mõlemad töövoo moodulid on avatud lähekoodiga ning kasutavad tarkvarakonteinereid, mis võimaldab kasutajatel neid lihtsalt laiendada ning jooksutada erinevates arvutuskeskkondades. Kvaliteedikontrolli teostamiseks ning andmete normaliseerimiseks arendasime välja skripti, mis automaatselt arvutab välja erinevad kvaliteedimõõdikud ning esitab need kasutajale. Juhtprojekti raames viisime läbi geeniekspressiooni QTL analüüsi 15 andmestikus ja 40 erinevas bioloogilises kontekstis ning tuvastasime vähemalt ühe statistiliselt olulise QTLi enam kui 9000 geenile. Loodud töövoogude laialdasem kasutuselevõtt võimaldab muuta QTL analüüsi korratavamaks, teisaldatavamaks ning lihtsamini kasutatavaks.Quantitative trait locus (QTL) analysis links variations in molecular phenotype expression levels to genotype variation. This analysis has become a standard practice to better understand molecular mechanisms underlying complex traits and diseases. Typical QTL analysis consists of multiple steps. Although a diverse set of tools is available to perform these individual analysis, the tools have so far not been integrated into a reproducible and scalable workflow that is easy to use across a wide range computational environments. Our analysis workflow consists of three modules. The analysis starts with quantification of the phenotype of interest, proceeds with normalisation and quality control and finishes with the QTL analysis. For phenotype quantification and QTL mapping modules we developed pipelines following best practices of the nf-core framework. The pipelines are containerized, open-source, extensible and eligible to be parallelly executed in a variety computational environments. For quality control module we developed a script which automatically computes the measures of quality and provides user with information. As a proof of concept, we uniformly processed more than 40 context specific groups from more than 15 studies and discovered at least one significant eQTL for more than 9000 genes. We believe that adopting our pipelines will increase reproducibility, portability and robustness of QTL analysis in comparison to existing approaches

    Processed and annotated yeast gene expression data from yeast2 and ygs98 platforms

    No full text
    <p>This dataset contains the following files:</p> <ul> <li><em>yeast2_processed_rds.tar.gz -</em> processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>ygs98_processed_rds.tar.gz </em>- processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>yeast2-curated-annotations.txt</em> - metadata for the yeast2 platform.</li> <li><em>ygs98-curated-annotations.txt</em> - metadata for the ygs98 platform.</li> </ul> <p> </p&gt

    Chromatin accessibility QTL lead variants in macrophages stimulated with IFNg and Salmonella

    No full text
    <p>Lead caQTL variants from RASQUAL and FastQTL analyses.</p&gt

    Predicting the Impact of Non-Coding Genetic Variants on Transcription Factor Binding with Machine Learning

    Get PDF
    Inimorganismi toimimispõhimõtetest arusaamine on üks tänapäeva teaduse suurimaid väljakutseid. Esimese inimgenoomi sekveneerimise järgselt on palju resursse kulutatud DNA sekventsi ja selle varieeruvuse uurimiseks. Nendest jõupingutustest hoolimata ei saa me paljudest inimorganismis toimuvatest olulistes protsessidest endiselt väga hästi aru. Üks selliseid protsesse on mittekodeerivate geneetiliste variantide mõju hindamine inimese genoomis. Sellised geneetilised variandid, kui neil üldse peaks mingi mõju olema, mõjutavad suure tõenäosusega transkriptsioonifaktorite seondumist. Suur hulk erinevaid meetodeid on välja töötatud ennustamaks geneetiliste variantide mõju transkriptsioonifaktorite seondumisele. Täpsete testandmestike puudumise tõttu on aga nende meetodite täpsuse hindamine olnud raskendatud ja seetõttu on enamasti lähtutud kaudsetest mõõdikutest. Oma töös panen ma kokku kolm suurt geneetilist andmestikku, et kindlaks teha suur hulk geneetilisi variante, mis suure tõenäosusega mõjutavad põhjuslikult kahe transkriptsioonifaktori (CTCF ja PU.1) seondumist DNAle. Järgnevalt kasutan ma neid geneetilisi variante hindamaks kolme kaasaegse ennustusalgoritmi täpsust. Minu tulemused näitavad, et kuigi mõne suure mõjuga geneetilise variandi efekti hindamine on võimalik, jääb enamiku väiksema mõjuga variantide mõju kindlaks tegemata. See lähenemine on üldistatav teistele transkriptsioonifaktoritele ja seda saab kasutada uudsete ennustusalgoritmide täpsuse paremaks hindamiseks tulevikus.Understanding how the human organism works is one of the most important problems in the science. A lot of research effort went into analysis of deoxyribonucleic acid (DNA) since the first human genome was sequenced. Despite these efforts, there are still a large number of poorly understood processes happening in the human organism. One of them is understanding the functional consequences of non-coding genetic variants in the DNA sequence of a human. These variants, if functional, are likely to influence the binding of transcription factors - regulatory proteins that control the expression of other genes by binding to regulatory elements across the genome. A diverse set of methods have been developed to predict the effect of genetic variants on transcription factor binding. However, all of these methods have been limited by the lack of high quality testing data to evaluate their accuracy. Here I combine and re-analyse three large genetic studies to identify a high quality set of likely causal genetic variants that regulate the binding of CTCF and PU.1 transcription factors. I then use these variants to evaluate the accuracy of three state-of-the art prediction algorithms. My results indicate that while the impact of some genetic variants with large effect can be readily predicted, most variants with smaller effects are missed by current prediction algorithms. My approach is generalisable to other transcription factors and can be used to benchmark the accuracy of novel prediction algorithms developed in the future

    Raw microarray gene expression datasets included in the eQTL Catalogue

    No full text
    Raw microarray intensity values for five datasets: CEDAR Fairfax_2012 Fairfax_2014 Naranbhai_2015 Kasela_2017 </ul

    Datasets associated with figures from the main paper

    No full text
    &lt;p&gt;This dataset contains data&nbsp;objects associated with the figures in the paper &quot;Shared genetic effects on chromatin and gene expression indicate a role for enhancer priming in immune response&quot; (https://doi.org/10.1038/s41588-018-0046-7).&lt;/p&gt

    Sample paired-end RNA-seq data from four human macophage samples (chromsome 21 only)

    No full text
    &lt;p&gt;This is a small subset of the RNA-seq data from a recently published macrophage study (https://doi.org/10.1038/s41588-018-0046-7) intended for quickly testing alignment and quantification methods. The fastq files only contain reads mapping to the chromosome 21 of the human genome.&lt;/p&gt; &lt;p&gt;The dataset consists of macrophages from two individuals (eipl and fikt) profiled in two conditions:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;A - naive&lt;/li&gt; &lt;li&gt;C - 5 hours after Salmonella infection&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;The data is paired end with 1 and 2 denoting the first and second pair, respectively.&lt;/p&gt

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Sample paired-end RNA-seq data from four human macophage samples (chromsome 21 only)

    No full text
    This is a small subset of the RNA-seq data from a recently published macrophage study (https://doi.org/10.1038/s41588-018-0046-7) intended for quickly testing alignment and quantification methods. The fastq files only contain reads mapping to the chromosome 21 of the human genome. The dataset consists of macrophages from two individuals (eipl and fikt) profiled in two conditions: A - naive C - 5 hours after Salmonella infection The data is paired end with 1 and 2 denoting the first and second pair, respectively
    corecore