1,721,030 research outputs found

    Inferring Developmental Trajectories and Causal Regulations with Single-cell Genomics

    Get PDF
    Thesis (Ph.D.)--University of Washington, 2018Development is commonly regarded as a hierarchical branching process which is governed by underlying gene regulatory networks. Single-cell genomics, single-cell RNA-seq (scRNA-seq) in particular, holds the promise to resolve the dynamics of this process. However, learning the structure of complex single-cell trajectories with multiple branches remains a challenging computational problem. In this thesis, I will present the toolkit, Monocle 2, which uses reversed graph embedding to reconstruct single-cell trajectories in a fully unsupervised manner. Monocle 2 learns an explicit “principal graph” that passes through the middle of the data as opposed to other ad hoc methods, greatly improving the robustness and accuracy of its trajectories. I will demonstrate that Monocle 2 is able to accurately reconstruct developmental trajectories for complicated systems, including hematopoiesis involving multiple different cell fates. When coupled with another statistical framework, BEAM (branch expression analysis modeling), Monocle 2 is able to detect genes specific to different developmental lineages. The unprecedented high resolution of the reconstructed developmental trajectories not only enables us to determine which genes are playing important roles at the critical time point of cell fate transition but also to directly infer causal gene regulatory networks. To this end, I have been developing a new toolkit, Scribe, which applies novel information theory techniques to detect causal interactions responsible for fate transitions. Scribe provides intuitive visualizations of causal interactions and can additionally incorporate information from “RNA-velocity” for causality detection. Scribe accurately reconstructs core networks specifying myelocytic or chromaffin cells. Finally, I will show a compendium of the inferred causal regulatory network for C elegans’ early embryogenesis based on lineage resolved live imaging data, demonstrating Scribe’s generalizability

    Development and application of combinatorial single cell methods to tissue physiology and disease

    Get PDF
    Thesis (Ph.D.)--University of Washington, 2022Advances in single-cell sequencing technologies present an opportunity to study the cellular and molecular basis of complex tissue physiology at global scale with unprecedented throughput. In this work, methods for application of a method of single cell sequencing – single-cell combinatorial indexing (“sci”) methods – were extended for use in mammalian tissue. In the human heart, analysis of single nucleus data reveals variation by age and sex in healthy donors, divergent use of transcription factor motifs in adult vs. fetal chromatin, and distal genomic sites that improve predictive models of cell type-specific expression. In mouse models of lung disease, we find aberrant differentiation states in the macrophages of a model of pulmonary alveolar proteinosis, emergence of an osteoclast-like macrophage phenotype in a model of silicosis, and in both models we find widespread alterations across the cell types of the lung. In overlapping work, we build interpretable predictive models of gene expression in the human heart and in the P. falciparum parasite, finding cell-type specific transcription factors in adult human heart and uncovering strikingly stage-specific information content in histone marks in Plasmodium

    Identifying Neural Pathways of Stress and Fear with Retrograde Viral Tracing and Single Cell RNA Sequencing

    No full text
    Thesis (Ph.D.)--University of Washington, 2022The brain is made up of millions of neurons, often making thousands of connections each. Interrogating individual neurons activated by stressors is the equivalent of finding a needle in a haystack. Here we describe experimental and computational methods for tracing neural connections and identifying neurons involved in the stress-response pathway of corticotropin releasing hormone (CRH) signaling. We perform single cell RNA sequencing on neurons upstream of CHR neurons (CRHNs) and show incredible diversity and co-expression of signaling molecules. Subsequent studies sort cells to enrich for activated neurons upstream of CHRNs in response to restraint stress. This analysis reveals known neuron subtypes in activated CHRNs and reports new genes of interest in CHRN signaling in response to restraint stress. These studies provide a framework for mapping neural connections upstream of targeted cell populations, and the identification of activated neurons through computational analysis

    Transcriptomic profiling of macrophage polarization

    No full text
    Thesis (Ph.D.)--University of Washington, 2020Macrophages perform a wide variety of crucial, and sometimes contradictory, functions. While “pro-inflammatory” activities like fighting off infections and “anti-inflammatory” activities like wound-healing traditionally have been attributed to M1 and M2 macrophage subsets, studies of macrophages in both in vitro and in vivo contexts suggest that macrophages are phenotypically plastic and may shift states in response to environmental changes. However, it remains unclear whether macrophages retain any persistent memory of past polarization states which may then impact their future repolarization to new states. In this dissertation, I first describe the evolving understanding of macrophage polarization and phenotypic plasticity and introduce commonly used models for macrophage polarization in humans and mice. I also outline recent advances in single-cell RNA-sequencing and some of the most popularly used platforms for single-cell transcriptomics. I then focus on my work characterizing macrophage polarization and repolarization in vitro, where I performed deep transcriptomic profiling at high temporal resolution as macrophages were polarized with cytokines that drive them into “M1” and “M2” molecular states. I find through trajectory analysis of their global transcriptomic profiles that macrophages which are first polarized to M1 or M2 and then subsequently repolarized demonstrate little to no memory of their polarization history. I observe complete repolarization both from M1 to M2 and vice versa, and I find that macrophage transcriptional phenotypes are defined by the current cell microenvironment, rather than an amalgamation of past and present states. In the following chapters, I present preliminary work aimed at characterizing alveolar macrophages, the tissue-resident macrophages of the lung: I first describe my attempts to identify key stimuli for triggering specification of alveolar macrophage fate. I then discuss preliminary results from single-cell RNA-seq profiling of alveolar macrophages in the context of pulmonary alveolar proteinosis. In the final chapter, I reflect on challenges I faced in adapting single-cell RNA-sequencing methods to work with primary tissue cells and summarize the main findings from my work

    Algorithms for modeling gene regulation and determining cell type using single-cell molecular profiles

    Get PDF
    Thesis (Ph.D.)--University of Washington, 2019Single-cell genomic technologies are helping us answer key biological questions that have long remained elusive. How do a single cell and a single genome generate such complex multicellular organisms as humans? More specifically, how do these cells orchestrate specific transcriptional programs depending on their cell type? New technologies like single-cell RNA-seq and single-cell ATAC-seq allow us to examine the transcription and regulation of individual cells as they develop; however, these methods have important limitations. A primary limitation with all single-cell data is data sparsity, which must be overcome computationally to extract useful information from these experiments. In this dissertation, I present two algorithms designed to overcome the sparsity of single-cell data and allow biological discovery. I first introduce Cicero for single-cell chromatin accessibility data, which is both an algorithm that calculates co-accessibility scores to assign distal regulatory elements to genes, and a software system that adapts existing single-cell RNA-seq analysis techniques for use with single-cell chromatin accessibility data. In Chapter 2, I apply Cicero to an in vitro myoblast differentiation assay and find evidence for the use of ”chromatin hubs” during myogenesis. In Chapter 3, I apply Cicero to single-cell ATAC-seq data from mouse bone marrow and recapitulate known patterns of hematopoiesis and known cis-regulation of the b-globin locus. In Chapter 4, I introduce a second algorithm, Garnett, which uses single-cell expression data to train and apply automated cell type classifiers. The accuracy of this technology is demonstrated with data from various single-cell RNA-seq methods and tissue sources. In a final chapter, I reflect on the development of software for biological applications and future directions for this work

    A molecular atlas of C. elegans development at single-cell and single-lineage resolution

    No full text
    Thesis (Ph.D.)--University of Washington, 2019It takes many cell divisions to produce a complex, multicellular organism such as a human being. Every one of the trillions of cells in a human body was produced by the division of a parent cell, which in turn was produced by the division of its parent; and if one follows this lineage far back enough, one will reach the single zygote cell that is the common progenitor of all of the cells in the body. Collectively, the pattern of cell divisions that produce an organism is called its cell lineage. As cells divide in a developing organism, they also differentiate into specialized cell types. What cell type a cell will adopt—its “cell fate”—is restricted by its lineage. A cell that descends from an endoderm progenitor, for example, may differentiate into a liver cell or an intestine cell, but will not become a bone cell. This general principle, that different parts of an organism’s cell lineage have different developmental potentials, has been known since the early 1800s. But our understanding of the molecular mechanisms that connect the cell lineage to the process of cell differentiation remains incomplete. In this dissertation, I present a near-comprehensive atlas of gene expression in the embryonic cell lineage of the nematode Caenorhabditis elegans, the only animal for which the cell lineage is fully known. I describe the methods used to assemble this atlas from single cell RNA-seq data, which required finding the precise lineage identity of each assayed cell. Using the atlas, I investigate the molecular mechanisms of cell fate commitment, finding that: 1. Multilineage priming is strikingly prevalent, contributing to the differentiation of over half of the cells in the lineage. 2. Distinct lineages that produce the same anatomical cell type tend to converge to a homogenous transcriptional state. This convergence is gradual for lineages that commit to their cell fate early in development, but can be abrupt for lineages that commit late. 3. A cell’s lineage and its transcriptome are correlated, but this correlation is transient, peaking in late gastrulation and falling dramatically during terminal differentiation. 4. Developmental trajectories reconstructed from single cell RNA-seq data often do not accurately reflect the cell lineage, in large part due to the transcriptional convergence of similarly fated lineages. Supplementary datasets, e.g. from fluorescent reporter imaging, are necessary to accurately place single cell RNA-seq data in the context of the cell lineage. This work provides an extensive resource to the C. elegans research community and an outline of the challenges that will need to be overcome in future studies of vertebrate cell lineages by single cell RNA-seq

    Embryo-scale single-cell chemical transcriptomics reveals dependencies between cell types and signaling pathways

    No full text
    Thesis (Ph.D.)--University of Washington, 2024Organogenesis is a fast and coordinated process that is highly conserved across vertebrates. However, we lack a detailed map of how each developing cell type in each organ depends on the developmental signaling pathways. Conventional gene knockouts cannot target an entire signaling pathway or have severe early phenotypes, making it challenging to study regulation of organogenesis. To fill this gap in knowledge, we applied highly multiplexed, scalable and temporally controlled in vivo chemical perturbations to dissect individual signaling pathways’ (BMP, FGF, Notch, RA, Shh, TGFβ & Wnt) roles in organogenesis. Here we present the transcriptional data of 711,997 cells from 288 chemically perturbed zebrafish embryos, spanning multiple time points (36, 48 and 72 hours post fertilization (hpf)) in organogenesis. The high degree of replication in our approach, which independently profiles multiple embryos per condition with DNA barcodes using Sci-Plex, allows us to statistically calculate the variance in cell type abundance and gene expression, embryo-wide and detect perturbation-dependent changes, even in very rare cell types. With this approach, I define a list of cell types that utilize each signaling pathway as well as undergo cross-regulation from other signaling pathways. I demonstrate that differential global gene expression analysis can identify novel signaling pathway regulation of transcription factors (TFs) in select cell types. I characterize novel signaling pathway regulation of multiple cell types within the pectoral fin mesoderm, a very rare tissue comprising less than 1% of the embryo. Shh is known to positively regulate pectoral fin mesoderm and cartilage, however, we show Shh also negatively regulates cleithrum size. In addition to comprehensively characterizing signaling pathway regulation of the pectoral fin over time, I constructed the highest-resolution trajectory to date of the developing pectoral fin mesoderm and identified two previously uncharacterized lateral plate mesoderm (LPM)-derived cell types as well as multiple novel gene markers of these LPM-derived pectoral fin cell types. Transcriptional and spatial analysis characterized one of the unknown cell types as distal mesenchyme, which was previously mischaracterized as an ectoderm-derivative, and the other unknown cell type as a tenocyte population that we show lie in between the muscle and the cartilage, presumably serving a functional role of anchoring the two cell types together. Taken together, our data provides a valuable resource for understanding the development of paired appendages and the comprehensive roles of signaling pathways in regulating vertebrate organogenesis in embryos as a whole.In this dissertation, I will first introduce how essential signaling pathways are in regulating embryonic development and how previous techniques to study their roles were neither comprehensive nor high-throughput, limiting discovery of novel biology. I will introduce how single-cell technologies can fill the need to profile embryos in an unbiased, comprehensive and high-throughput way. Using zebrafish pectoral fins as an example, I show how we can use this approach to develop models of signaling pathway regulation within individual tissue lineages

    Algorithms for differential analysis of cellular composition in single-cell perturbation experiments

    No full text
    Thesis (Ph.D.)--University of Washington, 2024Advancements in multiplexing techniques have enabled the application of single-cell genomic methods to comprehensively study the effects of high-throughput perturbation experiments at a whole-embryo scale. Such analyses aim to pinpoint key genes, cell types, and signaling pathways that control cell fate decisions during development. However, there is a lack of statistically principled tools for measuring how cell types shift after perturbations (genetic, chemical, or environmental) and identifying which genes regulate those transitions. In this thesis, I introduce two new software packages for studying single-cell perturbation experiments. Hooke is a new software package that uses Poisson-Lognormal models to perform differential analysis of cell abundances for perturbation experiments read out by single-cell RNA-seq. This versatile framework allows users to 1) perform multivariate statistical regression to describe how perturbations alter the relative abundances of each cell state and 2) describe how all pairs of states co-vary as a parsimonious network of partial correlations. To demonstrate Hooke’s utility, we analyzed a single-cell atlas of zebrafish organogenesis that includes wild-type and genetic perturbations at whole-embryo scale across multiple time points. This method identified novel genetic requirements for relatively rare cell types in the embryonic kidney. Platt is another new software package that uses Hooke's outputs to construct lineage graphs based on the covariation of cell type counts in time series and perturbation data. These graphs help identify candidate transcription factors important in lineage specification and organize differential abundance results into direct and indirect effects. With Platt, we study the impact of knocking out the lmx1b, a homeobox transcription factor with specific expression in multiple lineages. Both packages aim to fill a critical gap by allowing users to characterize how their experimental perturbations alter cells' proportions and molecular states in complex tissues or whole embryos

    Expanding the scope and utility of single-cell genomic technologies

    No full text
    Thesis (Ph.D.)--University of Washington, 2019Technological advancements in single-cell genomic technologies have led to an exponential increase in number of cells from which biologists are able to measure various molecular profiles. This dramatic increase in scalability opens up the potential to leverage such technologies as the basis for a wide array of off-label applications. For example, the genetic screening field has typically been limited approaches that examine changes in relative abundance of genotypes represented in a given cell population before and after screening/selection as a proxy for phenotype. Molecular measurements such as RNA sequencing (RNA-seq) could allow a more generic and direct readout of molecular phenotypes, but performing bulk RNA-seq on many mutants in an arrayed format dramatically limits scalability. Single-cell RNA-seq, if modified to enable the readout of cell genotypes, would allow us to examine the molecular changes resulting from individual genotypes in a pooled format. In fact, single-cell technologies create the tremendous opportunity of multi-modal measurements (e.g. gene expression in addition to genotypes, lineage information, protein abundances via oligo tagged antibodies, T and B cell receptor sequences, entire additional molecular measurements like chromatin accessibility). An important step in the progression towards this goal is the adaptation of other molecular profiling techniques beyond RNA sequencing to single-cell formats, which come with their own experimental and analytical challenges that have yet to be solved within the field. In this dissertation, I first introduce an attempt to use single-cell transcriptomics to serve as a readout for genetic screens. In Chapter 2, I introduce our work to enable such a readout from CRISPR and CRISPRi-based genetic screens and find that a number of approaches that have been used to tackle this problem suffer from high error rates (~50\%) in the assignment of correct genotypes to cells. Ultimately we recommend a particular existing design that does not suffer from this common flaw. In Chapter 3, we utilize this method in conjunction with highly multiplexed perturbations to enable the screening of almost 6000 putative regulatory elements in the genome, measuring their impact on gene expression in cis. Notably an experiment of this scale would have been extremely difficult to achieve without the use of single-cell technologies. Lastly, in Chapter 4, I shift focus to our efforts to expand the set of molecular measurements that we can make with single-cell technologies. Specifically, I describe our efforts to measure chromatin accessibility using single-cell ATAC-seq across 13 different tissues in 8-week old mice and how we went about addressing the many computational challenges that arise when attempting to analyze and interpret such datasets

    Novel single-cell genomic approaches for deciphering cellular heterogeneity

    No full text
    Thesis (Ph.D.)--University of Washington, 2025Single-cell genomics has reshaped our understanding of developmental biology by uncovering intricate molecular states at unprecedented scale. However, continued development of experimental and computational approaches is needed to fully realize its potential. In this dissertation, I will introduce three distinct projects, each centered around adapting computational tools or developing novel experimental approaches to decipher cellular heterogeneity. In the first project, we adapted latent Dirichlet allocation to single-cell combinatorial indexed Hi-C (sci-Hi-C) intra-chromosomal contact maps to decompose the data into chromatin topics. Our approach enabled co-embedding and clustering of sci-Hi-C data derived from five different cell lines (GM12878, H1Esc, HFF, IMR90, and HAP1) and identification of cell type-specific topics of chromatin interactions. In the second project, we developed inexpensive spike-in controls for single-cell combinatorial indexed RNA-seq (sci-RNA-seq) experiments using a set of single-stranded hash oligonucleotides (“hash ladder”). To normalize for technical variation introduced within individual cells, we calculated a cell-specific size factor that is derived from the hash ladder. We applied the ladder to study the effects of various chemical perturbations, including RNA pol II elongation, histone deacetylation, and activation of the glucocorticoid receptor. In the third project, we vastly improved Visual Cell Sorting (VCS), which is an automated imaging workflow that enables binning and sorting of cells by visual phenotypes, by making it compatible with cell fixation and three-level sci-RNA-seq (sci-RNA-seq3). We applied VCS to sort over one million E15 F1 B6xCAST mouse embryo derived nuclei based on nucleolar and nuclear speckle size, and we profiled the sorted nuclei with sci-RNA-seq3. We revealed differences in these nuclear compartment sizes within and across cell types and identified both expected and unexpected correlations with proliferation and differentiation status. We identified 42 genes that positively correlated with relative nucleolar size, reflecting the activation of gene expression programs relating to ribosomal biogenesis and proteostasis stress response. Finally, we demonstrated that these genes can be used to quantify relative nucleolar size across mouse, human, and zebrafish developmental atlases
    corecore