Cold Spring Harbor Laboratory

Cold Spring Harbor Laboratory Institutional Repository
Not a member yet
    13273 research outputs found

    A Correspondence between Normalization Strategies in Artificial and Biological Neural Networks.

    No full text
    A fundamental challenge at the interface of machine learning and neuroscience is to uncover computational principles that are shared between artificial and biological neural networks. In deep learning, normalization methods such as batch normalization, weight normalization, and their many variants help to stabilize hidden unit activity and accelerate network training, and these methods have been called one of the most important recent innovations for optimizing deep networks. In the brain, homeostatic plasticity represents a set of mechanisms that also stabilize and normalize network activity to lie within certain ranges, and these mechanisms are critical for maintaining normal brain function. In this article, we discuss parallels between artificial and biological normalization methods at four spatial scales: normalization of a single neuron's activity, normalization of synaptic weights of a neuron, normalization of a layer of neurons, and normalization of a network of neurons. We argue that both types of methods are functionally equivalent-that is, both push activation patterns of hidden units toward a homeostatic state, where all neurons are equally used-and we argue that such representations can improve coding capacity, discrimination, and regularization. As a proof of concept, we develop an algorithm, inspired by a neural normalization technique called synaptic scaling, and show that this algorithm performs competitively against existing normalization methods on several data sets. Overall, we hope this bidirectional connection will inspire neuroscientists and machine learners in three ways: to uncover new normalization algorithms based on established neurobiological principles; to help quantify the trade-offs of different homeostatic plasticity mechanisms used in the brain; and to offer insights about how stability may not hinder, but may actually promote, plasticity

    Cross-tissue analysis of allelic X-chromosome inactivation ratios resolves features of human development

    Get PDF
    X-chromosome inactivation (XCI) is a random, permanent, and developmentally early epigenetic event that occurs during mammalian embryogenesis. We harness these features of XCI to investigate characteristics of early lineage specification events during human development. We initially assess the consistency of X-inactivation and establish a robust set of XCI-escape genes. By analyzing variance in XCI ratios across tissues and individuals, we find that XCI is completed prior to tissue specification and at a time when 6-16 cells are fated for all tissue lineages. Additionally, we exploit tissue specific variability to characterize the number of cells present at the time of each tissue’s lineage commitment, ranging from approximately 20 cells in liver and whole blood tissues to 80 cells in brain tissues. By investigating variance of XCI ratios using adult tissue, we resolve key features of human development otherwise difficult to ascertain experimentally and develop scalable methods easily applicable to future data

    High resolution copy number inference in cancer using short-molecule nanopore sequencing.

    Get PDF
    Genome copy number is an important source of genetic variation in health and disease. In cancer, Copy Number Alterations (CNAs) can be inferred from short-read sequencing data, enabling genomics-based precision oncology. Emerging Nanopore sequencing technologies offer the potential for broader clinical utility, for example in smaller hospitals, due to lower instrument cost, higher portability, and ease of use. Nonetheless, Nanopore sequencing devices are limited in the number of retrievable sequencing reads/molecules compared to short-read sequencing platforms, limiting CNA inference accuracy. To address this limitation, we targeted the sequencing of short-length DNA molecules loaded at optimized concentration in an effort to increase sequence read/molecule yield from a single nanopore run. We show that short-molecule nanopore sequencing reproducibly returns high read counts and allows high quality CNA inference. We demonstrate the clinical relevance of this approach by accurately inferring CNAs in acute myeloid leukemia samples. The data shows that, compared to traditional approaches such as chromosome analysis/cytogenetics, short molecule nanopore sequencing returns more sensitive, accurate copy number information in a cost effective and expeditious manner, including for multiplex samples. Our results provide a framework for short-molecule nanopore sequencing with applications in research and medicine, which includes but is not limited to, CNAs

    The clinical behavior and genomic features of the so-called adenoid cystic carcinomas of the solid variant with basaloid features.

    No full text
    Classic adenoid cystic carcinomas (C-AdCCs) of the breast are rare, relatively indolent forms of triple negative cancers, characterized by recurrent MYB or MYBL1 genetic alterations. Solid and basaloid adenoid cystic carcinoma (SB-AdCC) is considered a rare variant of AdCC yet to be fully characterized. Here, we sought to determine the clinical behavior and repertoire of genetic alterations of SB-AdCCs. Clinicopathologic data were collected on a cohort of 104 breast AdCCs (75 C-AdCCs and 29 SB-AdCCs). MYB expression was assessed by immunohistochemistry and MYB-NFIB and MYBL1 gene rearrangements were investigated by fluorescent in-situ hybridization. AdCCs lacking MYB/MYBL1 rearrangements were subjected to RNA-sequencing. Targeted sequencing data were available for 9 cases. The invasive disease-free survival (IDFS) and overall survival (OS) were assessed in C-AdCC and SB-AdCC. SB-AdCCs have higher histologic grade, and more frequent nodal and distant metastases than C-AdCCs. MYB/MYBL1 rearrangements were significantly less frequent in SB-AdCC than C-AdCC (3/14, 21% vs 17/20, 85% P < 0.05), despite the frequent MYB expression (9/14, 64%). In SB-AdCCs lacking MYB rearrangements, CREBBP, KMT2C, and NOTCH1 alterations were observed in 2 of 4 cases. SB-AdCCs displayed a shorter IDFS than C-AdCCs (46.5 vs 151.8 months, respectively, P < 0.001), independent of stage. In summary, SB-AdCCs are a molecularly heterogeneous but clinically aggressive group of tumors. Less than 25% of SB-AdCCs display the genomic features of C-AdCC. Defining whether these tumors represent a single entity or a collection of different cancer types with a similar basaloid histologic appearance is warranted

    Deconvolution of Expression for Nascent RNA Sequencing Data (DENR) Highlights Pre-RNA Isoform Diversity in Human Cells

    Get PDF
    Quantification of mature-RNA isoform abundance from RNA-seq data has been extensively studied, but much less attention has been devoted to quantifying the abundance of distinct precursor RNAs based on nascent RNA sequencing data. Here we address this problem with a new computational method called Deconvolution of Expression for Nascent RNA sequencing data (DENR). DENR models the nascent RNA read counts at each locus as a mixture of user-provided isoforms. The performance of the baseline algorithm is enhanced by the use of machine-learning predictions of transcription start sites (TSSs) and an adjustment for the typical “shape profile” of read counts along a transcription unit. We show using simulated data that DENR clearly outperforms simple read-count-based methods for estimating the abundances of both whole genes and isoforms. By applying DENR to previously published PRO-seq data from K562 and CD4+ T cells, we find that transcription of multiple isoforms per gene is widespread, and the dominant isoform frequently makes use of an internal TSS. We also identify > 200 genes whose dominant isoforms make use of different TSSs in these two cell types. Finally, we apply DENR and StringTie to newly generated PRO-seq and RNA-seq data, respectively, for human CD4+ T cells and CD14+ monocytes, and show that entropy at the pre-RNA level makes a disproportionate contribution to overall isoform diversity, especially across cell types. Altogether, DENR is the first computational tool to enable abundance quantification of pre-RNA isoforms based on nascent RNA sequencing data, and it reveals high levels of pre-RNA isoform diversity in human cells

    An anchored chromosome-scale genome assembly of spinach improves annotation and reveals extensive gene rearrangements in euasterids.

    Get PDF
    Spinach (Spinacia oleracea L.) is a member of the Caryophyllales family, a basal eudicot asterid that consists of sugar beet (Beta vulgaris L. subsp. vulgaris), quinoa (Chenopodium quinoa Willd.), and amaranth (Amaranthus hypochondriacus L.). With the introduction of baby leaf types, spinach has become a staple food in many homes. Production issues focus on yield, nitrogen-use efficiency and resistance to downy mildew (Peronospora effusa). Although genomes are available for the above species, a chromosome-level assembly exists only for quinoa, allowing for proper annotation and structural analyses to enhance crop improvement. We independently assembled and annotated genomes of the cultivar Viroflay using short-read strategy (Illumina) and long-read strategies (Pacific Biosciences) to develop a chromosome-level, genetically anchored assembly for spinach. Scaffold N50 for the Illumina assembly was 389 kb, whereas that for Pacific BioSciences was 4.43 Mb, representing 911 Mb (93% of the genome) in 221 scaffolds, 80% of which are anchored and oriented on a sequence-based genetic map, also described within this work. The two assemblies were 99.5% collinear. Independent annotation of the two assemblies with the same comprehensive transcriptome dataset show that the quality of the assembly directly affects the annotation with significantly more genes predicted (26,862 vs. 34,877) in the long-read assembly. Analysis of resistance genes confirms a bias in resistant gene motifs more typical of monocots. Evolutionary analysis indicates that Spinacia is a paleohexaploid with a whole-genome triplication followed by extensive gene rearrangements identified in this work. Diversity analysis of 75 lines indicate that variation in genes is ample for hypothesis-driven, genomic-assisted breeding enabled by this work

    Pan-genomic Matching Statistics for Targeted Nanopore Sequencing

    Get PDF
    Nanopore sequencing is an increasingly powerful tool for genomics. Recently, computational advances have allowed nanopores to sequence in a targeted fashion; as the sequencer emits data, software can analyze the data in real time and signal the sequencer to eject “nontarget” DNA molecules. We present a novel method called SPUMONI, which enables rapid and accurate targeted sequencing using efficient pan-genome indexes. SPUMONI uses a compressed index to rapidly generate exact or approximate matching statistics in a streaming fashion. When used to target a specific strain in a mock community, SPUMONI has similar accuracy as minimap2 when both are run against an index containing many strains per species. However SPUMONI is 12 times faster than minimap2. SPUMONI's index and peak memory footprint are also 16 to 4 times smaller than those of minimap2, respectively. This could enable accurate targeted sequencing even when the targeted strains have not necessarily been sequenced or assembled previously

    Using focal cooling to link neural dynamics and behavior.

    No full text
    Establishing a causal link between neural function and behavioral output has remained a challenging problem. Commonly used perturbation techniques enable unprecedented control over intrinsic activity patterns and can effectively identify crucial circuit elements important for specific behaviors. However, these approaches may severely disrupt activity, precluding an investigation into the behavioral relevance of moment-to-moment neural dynamics within a specified brain region. Here we discuss the application of mild focal cooling to slow down intrinsic neural circuit activity while preserving its overall structure. Using network modeling and examples from multiple species, we highlight the power and versatility of focal cooling for understanding how neural dynamics control behavior and argue for its wider adoption within the systems neuroscience community

    2,796

    full texts

    13,273

    metadata records
    Updated in last 30 days.
    Cold Spring Harbor Laboratory Institutional Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇