1,720,968 research outputs found
Computational Biology in Acute Myeloid Leukemia with CEBPA Abnormalities
__Abstract__
In the last decade, tiling-array and next-generation sequencing technologies allowed quantitative measurements of different cellular processes, such as mRNA expression, genomic changes including deletions or amplifications, DNA-methylation, chromatin modifications or Protein-DNA-binding interactions. Using these technologies, thousands of features can now be measured simultaneously in a patient cell sample. The use of for instance mRNA expression profiles or DNA-methylation profiles have already provided new insight into the molecular biology of patients with Acute Myeloid Leukemia (AML). AML is a blood cell malignancy, in which primitive myeloid cells have been transformed and accumulate in the bone marrow and blood. Different forms of AML exist with different molecular abnormalities that associate with distinct responses to therapy. Many subgroups with comparable mRNA expression or DNA-methylation patterns were identified. These studies also revealed the existence of novel previously undefined AML subtypes. Among those was a group of patients with a mutation in a gene called CEBPA. CEBPA is a gene that encodes the transcription factor CCAAT Enhancer Binding Protein Alpha (C/EBPα), which controls the expression of genes in myeloid progenitor cells. Mutated CEBPA encodes a dysfunctional C/EBPα-protein, which consequently results in aberrant control of “target genes”. In this thesis we focus particularly on the role of CEBPA. We studied the predictive and prognostic relevance of mutated CEBPA, and analyzed in a genome wide fashion the mRNA expression, DNA-methylation and the protein-DNA-binding levels corresponding to (mutated) CEBPA in AML. For the analysis of protein-DNA-binding, we developed a novel statistical methodology. With this statistical methodology we studied the fundamental role of (mutant) C/EBPα binding and the effect on gene expression levels. We also integrated gene expression with DNA-methylation profiles of hundreds of AML patients and revealed the existence of two previously unidentified AML subtypes
GTEx (Genotype-Tissue Expression) data normalized
This is a normalized dataset from the original RNAseq dataset downloaded from Genotype-Tissue Expression (GTEx) project: www.gtexportal.org: RNA-SeQCv1.1.8 gene rpkm Pilot V3 patch1.
The data was used to analyze how tissue samples are related to each other in terms of gene expression data The data can be used to get insights in how gene expression levels behave in in the different human tissues
2D Representation of Transcriptomes by t-SNE Exposes Relatedness between Human Tissues
The GTEx Consortium reported that hierarchical clustering of RNA profiles from 25 unique tissue types among 1641 individuals accurately distinguished the tissue types, but a multidimensional scaling failed to generate a 2D projection of the data that separates tissue-subtypes. In this study we show that a projection by t-Distributed Stochastic Neighbor Embedding is in line with the cluster analysis which allows a more detailed examination and visualization of human tissue relationships.Intelligent SystemsElectrical Engineering, Mathematics and Computer Scienc
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Hypergeometric analysis of tiling-array and sequence data: Detection and interpretation of peaks
Probing protein-deoxyribonucleic acid (DNA) is gaining popularity as it sheds light on molecular mechanisms that regulate the expression of genes. Currently, tiling-arrays and next-generation sequencing technology can be used to measure these interactions. Both methods generate a signal over the genome in which contiguous regions of peaks on the genome represent the presence of an interacting molecule. Many methods do exist to identify functional regions of interest (ROIs) on the genome. However the detection of ROIs are often not an end-point in research questions and it therefore requires data dragging between tools to relate the ROIs to information present in databases, such as gene-ontology, pathway information, or enrichment of certain genomic content. We introduce hypergeometric analysis of tiling-array and sequence data (HATSEQ), a powerful tool that accurately identifies functional ROIs on the genome where a genomic signal significantly deviates from the general genome-wide behavior. HATSEQ also includes a number of built-in post-analyses with which biological meaning can be attached to the detected ROIs in terms of gene pathways and de-novo motif analysis, and provides different visualizations and statistical summaries for the detected ROIs. In addition, HATSEQ has an intuitive graphic user interface that lowers the barrier for researchers to analyze their data without the need of scripting languages. We compared the results of HATSEQ against two other popular chromatin immunoprecipitation sequencing (ChIP-Seq) methods and observed overlap in the detected ROIs but HATSEQ is more specific in delineating the peak boundaries. We also discuss the versatility of HATSEQ by using a Signal Transducer and Activator of Transcription 1 (STAT1) ChIP-Seq data-set, and show that the detected ROIs are highly specific for the expected STAT1 binding motif.Intelligent SystemsElectrical Engineering, Mathematics and Computer Scienc
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
