1,720,979 research outputs found
Recommended from our members
Statistical algorithms in the study of mammalian DNA methylation
DNA methylation is a dynamic chemical modification that is abundant on DNA sequences and plays a central role in the regulatory mechanisms of cells. This modification can be inherited across cell divisions and generations, providing a ``memory mechanism" for regulatory programs that is more flexible than that coded in the DNA sequence. In recent years, high-throughput sequencing technologies have enabled genome-wide annotation of DNA methylation. Coupled with novel computational machinery, these developments have enabled unperceivable insight to the characteristics, biological function and disease association of this phenomenon. The collaborations between experimental and computational researches who take part in these efforts has been closer than ever before due to the need to involve computational methodologies throughout the entire research pipeline, from experimental design through bias correction to the analysis of large datasets. In the first part of this thesis we present contributions to the field of high-throughput DNA methylation. We introduce statistically sound criteria for the detection of methylation signatures in DNA sequence, and present an algorithm for the annotation of an informative non-overlapping subset of such regions that is optimal under biologically motivated assumptions. Our method outputs a sequence-generated list of regions that are of interest with respect to their methylation states. We then present a Bayesian network to infer corrected site-specific methylation states from a favorable but biased experimental method, and describe its incorporation in a software package. Along with site-specific methylation calls our package annotates experiment-specific regions of interest by considering both the methylation state inferences and the genomic sequence. These regions can serve as a basis for comparative methylation studies. In the last chapter of this section we bring results from a genome-scale comparative study conducted on humans, chimpanzees and an orangutan, providing evidence of DNA methylation differences that propagate through generations and distinguish these closely related species. The second part of this thesis concerns error correction in high-throughput sequencing datasets. In the course of studying DNA methylation with high-throughput sequencing we discovered a systematic error that results in false-positive variant detection and can significantly affect biological inferences in a variety of genomic studies. We present a classifier to correct for such errors and show that it performs very well with respect to both sensitivity and specificity
Recommended from our members
Methods and Applications of Differential Expression Analysis in Single-cell RNA-sequencing Data
Differential expression (DE) is one of the most commonly performed analyses when characterizing a single-cell RNA-seq (scRNA-seq) data set. This dissertation presents methods (Chapter 1 and 2) and applications (Chapter 3) of DE analysis in scRNA-seq data.
Chapter 1: Differential expression analysis, in spite of its wide application, is not trivial because the scRNA-seq data are high-dimensional, sparse, and noisy. In this chapter, we focus on the most simple but interesting setting: identifying genes with different mean expression levels between two groups of cells, assuming no complex dependencies among cells within each group. We proposed conditional differential expression, a framework for DE analysis, where we infer gene status of expressed (signal) or not (background), then apply DE algorithms and report results conditioning on the gene status. We discussed the interpretation of conditional DE results and showed that performance improvement could be achieved with a good gene status inference.
Chapter 2: The increasing accessibility of scRNAseq has encouraged the emergence of scRNAseq data from complex study designs, with batches over multiple biological replicates or diverse individuals. These data require an expanded definition of differential expression to capture the differences in the variabilities in addition to the means. In this chapter, we proposed a Poisson-lognormal multi-level model to account for both cell-to-cell and individual-to-individual variability. We provided two approaches to estimate the parameters: the method of moment estimators and the maximum likelihood estimator. Benchmarking against pseudobulk method confirmed that our model is not only useful in identifying changes in the mean expression levels across groups but also capable of capturing differences in the variance patterns, which could not be done by pseudobulk.
Chapter 3: In this chapter, we applied differential expression algorithms to a real data set of mouse Th17 cells, which are a subset of CD4 T cells that play an important role in autoimmunity. We analyzed Th17 cells collected from different tissues at homeostasis and/or during autoimmunity, characterizing their within and across tissues heterogeneity using unsupervised clustering followed by differential expression. In addition, we identified two subpopulations in the spleen during autoimmunity, inferred their migratory phenotype and plasticity using combined gene expression and T cell receptor information, and validated their functions with transfer and knockout experiments
Recommended from our members
Gene Expression Dynamics in Single Cells
Dynamic regulation of gene expression is central to fate choice and differentiation. Genes drive changes in cell state and also serve as markers of mature cell types and immature intermediates. Advances in single-cell RNA-seq (scSeq) have made it possible measure the expression of thousands of genes in tens of thousands of cells in a single experiment. Single-cell RNA sequencing (scSeq) opens up a new terrain in the study of differentiation and fate choice, but it has an important limitation: existing methods kill cells in the process of measurement, and thus prevent following a cell’s state over time. The inability to track cells prevents a full understanding of the dynamic events that accompany differentiation, or how variation in the state of progenitor cells predisposes their sub- sequent fate choices. These limitations are particularly evident in hematopoiesis – the process of steady-state blood production in bone marrow. scSeq is increasingly used to characterize the heterogeneity of hematopoietic stem and progenitor cells, but the functional consequences of this heterogeneity have been difficult to ascertain. Here, several methods are described for inferring single-cell gene expression dynamics. Chapter 1 presents a visualization method called SPRING that facilitates human inference of dynamics. A more direct approach to infer dynamics from high dimensional data is described in chapter 2 based on the principle of population balance. This analysis also highlights several properties of gene expression dynamics that cannot be inferred from a single-cell snapshot, including: whether unmeasured variables contribute significantly to fate choice; the degree to which cells in the same gene expression state also have the same velocity in their state; and where fate commitment occurs in gene expression space. Chapter 3 describes an experimental approach for tracking cell state over time that involves clonally marking cells, letting them divide, measuring one sister right away, and measuring the other sister later on. Applied to hematopoiesis, this method partially addresses the questions raised in Chapter 2, demonstrating the existence of hidden variables during differentiation and mapping fate probability across the gene expression landscape. It also opens up new directions for perturbational analysis of hematopoietic fate choice, and may pave the way for linking cell state and fate in other systems.Systems Biolog
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Recommended from our members
Uncovering the multi-faceted role of the tumor microenvironment with single-cell genomics: from identification of therapeutic targets to drivers of malignant cell states
Human cancer biology and response to therapy are driven by the diverse ecosystem of malignant and non-malignant cell types that compose a human tumor. Since the advent of single- cell genomics, this technique has been widely applied to study human malignancies and has highlighted the heterogeneity of the malignant cell population and the influence of the tumor microenvironment on malignant cell function and response to immunotherapy. These advances have deepened our understanding of the underlying biology, yet many malignancies remain difficult to treat. This dissertation will discuss two single-cell transcriptomics studies aimed at further understanding two such malignancies: primary brain tumors (gliomas) and Ewing sarcoma, a rare pediatric malignancy.
Chapter 1 presents a scRNA-seq study of glioma infiltrating T cells, where characterization of the T cell states identifies a novel therapeutic target present not only in gliomas but in other human malignancies. Chapter 2 focuses on the characterization of malignant cells in Ewing sarcoma and finds that a subset of malignant cell states are driven by interactions with the tumor microenvironment and predicts a worse prognosis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
