1,721,075 research outputs found
Identification of regulatory mechanisms in governing gene expression from the inactivated X chromosome
X chromosome inactivation (XCI) is the process in which one copy of the X chromosomes in females (XX) is randomly silenced to achieve dosage compensation of gene expression from X chromosome of males (XY). XCI is known to be incomplete, resulting in a subset of genes on the inactive X (Xi) being expressed. Genes along the X chromosome are classified into three categories based upon their expression from the Xi: subject, escape and variable. Escape genes, corresponding to 15% of X-linked genes, are generally expressed from the Xi at a substantially lower level compared to the active X (Xa). The underlying mechanisms in controlling escape from the heterochromatic state of Xi has been a long-standing question. While there is evidence that supports a role for intrinsic DNA elements in the escape XCI, they have not yet been identified.
With increasing amounts of data for transcription factor (TF) binding sites, the objective of this thesis was to identify the regulatory TFs facilitating the ability of genes to escape XCI via a bioinformatics approach using empirical data. ChIP-seq peaks of 155 TFs from the ReMap database were assessed for enrichment at regulatory regions of 55 escape genes. 19 TFs were identified via enrichment analysis in the transcription start site regions of escape genes. Co-binding of pairs of enriched TFs were characterized by gene set similarity between the target genes. Of the 155 TFs examined, ZFP36, an RNA binding protein that alters RNA stability, showed importance in both enrichment analysis and co-binding analysis.
An initial exploration of methods to compare the structures of Xi and Xa were undertaken using ChIA-PET data, providing insights into the limitations of current data resources and opportunities for future studies to inform the topographical organization of escape genes within Xi.
The results of this thesis refined our knowledge on cis-regulatory elements and trans-acting factors potentially involved in escape of XCI. This list of enriched TF may be useful in future analyses, experimental or computational, to further determine their sufficiency and necessity in the escape of XCI.Science, Faculty ofGraduat
Approaches to genome analysis through the application of graph theory
The human reference genome provides a framework against which the analysis and interpretation of an individual’s genome can be performed. Over the past twenty years the cost of genome sequencing has dropped from a prohibitive amount of hundreds of millions of dollars, to just a few thousand dollars. This has brought genome sequencing in line with the cost of other diagnostic medical tests, leading to a rapid uptake in both clinical and research settings. As a consequence of this global spread, deficiencies and population-specific inequities have emerged from the use of a framework that relies upon a single linear reference sequence. Partial, ad-hoc solutions, such as the introduction of alternative sequences for sections of the genome, have provided a stopgap but fail to fully represent the wealth of information now known about the level of variation that exists within and between populations. This thesis presents an alternative perspective on how we can take advantage of new computational methods to enhance the reference genome in the era of widespread sequencing and big data. An argument is given to motivate the revaluation of the role of the reference genome, and calls for a non-indexed, mutable reference framework with the crucial indexing methods to be shifted from the linear reference to a raw read set. A patented, edge-labelled, cyclic, graph-based model, the GNOmics Graph Model, is introduced as a flexible framework against which read alignment and variant calling can be performed. The value of indexing raw reads is explored through a published tool, FlexTyper, which allows a read set to be screened for informative markers. While there is still an ongoing global discussion as to how best to improve the reference genome, this thesis provides a thought-provoking reconceptualisation of applied human genome analysis.Science, Faculty ofGraduat
Advancing our understanding of genome regulation via optimization of stem cell differentiation and interpretable deep learning
The regulation of gene expression is a core challenge in understanding how diverse types of cells can be produced from the same DNA instructions. Insights about this complex machinery advance not only science but applications in therapy and pharmacology. For instance, the differentiation of stem cells for the purpose of regenerative medicine to treat patients with diabetes. In my second chapter, I address the problem of optimizing the differentiation protocol towards definitive endoderm, the precursor of insulin-producing pancreatic beta cells, by replacing the expensive growth factor with cheap molecule alternatives. I introduce a multiple-step pipeline based on small molecule transcriptome response profiles. The discovered chemicals emphasize the importance of key transcription factors in the process, such as HIF and MYC. The study of transcription factors is of high importance, and will further promote our knowledge about differentiation. Motivated by the thought, I explore the current trends of studying transcription factors in the gene regulation context. With large-scale data generation efforts by public consortia such as ENCODE, deep learning methods have become pervasive. A large training dataset is fundamental to the success of these methods, however, the amount of TF-related data is often small. To tackle this issue, in my third chapter, I perform an in-depth assessment of transfer learning for TF binding prediction and provide biologically motivated guidelines for efficient training of deep models when the data is limited. An additional challenge for deep models beyond data sufficiency is interpretability. In the fourth chapter, I systematically categorize and summarize interpretation approaches, exploring their underlying assumptions, strengths, and weaknesses. Inspired by transparent deep learning architectures, I present ExplaiNN, a new transparent model for the genomics tasks. I explore its efficiency and usability on a variety of problems in the fifth chapter of this thesis. Finally, in the last chapter, I apply ExplaiNN to ATAC-seq datasets of mouse and human immune systems to study differences in cis-regulatory logic. Transparency of the new method allowed me to discover a reproducible set of sequence motifs that either individually or combinatorially are responsible for the bulk of the predictions, and tend to have species-specific occurrence patterns.Science, Faculty ofGraduat
Identification of pharmacogenetic variants influencing the likelihood of developing cancer treatment-induced mucositis using pathway analyses
Methotrexate (MTX), a cornerstone cancer treatment, has contributed to improved 5-year event-free survival rates but is associated with 20-40% occurrence of mucositis, which is characterized by the development of painful inflammatory lesions largely focused along the alimentary tract that often leads to premature termination of cancer treatment and impacts survival. Previous studies identified individual genetic variants implicated in mucositis development. Considering the complex biological pathways involved in mucositis development highlighted in recent studies, we hypothesize that multiple genetic variants within these shared biological pathways contribute to the onset of MTX-induced mucositis.
To identify gene pathways that are likely to impact mucositis risk, we captured genes associated with: (i) methotrexate pharmacokinetics/pharmacodynamics from PharmGKB; (ii) published pathobiological pathways (e.g., WNT/β-catenin signaling) predicted to underlie mucositis development using MSigDB/Enrichr; and novel pathways based on genes previously associated with mucositis from literature using StringDB. To pinpoint pathways highly enriched for genetic variations associated with mucositis development in methotrexate-treated children, we examined the joint association of genetic variants using the raw genome-wide genotyping and patient clinical data for a set of pediatric mucositis cases (n=131) and controls (n=366) treated with intravenous methotrexate across 6 Canadian academic hospitals. A final set of 18 non-redundant pathways with a priori evidence for association with treatment-induced mucositis were selected for analysis. Our study used a phenotypic permutation test to identify significant enrichment in the IL-6 and WNT/β-catenin signaling pathways in patients developing mucositis due to high-dose MTX (> 1000 mg/m2). Using these genetic findings and clinical data, we developed a machine learning algorithm for mucositis risk stratification in patients receiving high-dose IV-MTX and identified key features impacting risk. Our future direction will focus on replicating the findings. If validated, the results can guide clinical care for high-risk patients and inform diagnosis and prescribing decisions. Moreover, the highlighted pathobiological pathways may serve as future therapeutic targets for mitigating mucositis during cancer treatment. Our goal is to minimize MTX-induced adverse reactions and improve patient quality of life during cancer treatment.Science, Faculty ofGraduat
Cell-conditional generative adversarial network
With single cell sequencing advances, research has increasingly focused on under-standing cell-specific gene regulation mechanisms. However, single cell sequencing data are often noisy and the amount of sequence obtained from rare cell types small. Simulation can be a powerful approach to aid understanding when data is limited, both because the process used to generate such data can provide mechanistic insights into cell-specific regulation and the data produced can augment analysis methods development. We constructed and optimized a stand-alone cell-conditional GAN (ccGAN) to simulate cell-specific ATAC-seq data. We trained our model on published single cell ATAC-seq (scATAC-seq) data that had been produced with different protocols on embryonic mice forebrain and adult mice brain. The ccGAN generated sequence was correlated in both Transcription Factor (TF) binding motif composition and positional distribution with the experimental scATAC-seq. The ccGAN simulator was able to learn important cell-specific signals amidst noise. The ccGAN architecture holds broad potential for single cell regulatory data simulation beyond ATAC-seq, such as for ChIP-seq or epigenetic propertiesScience, Faculty ofGraduat
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
