1,720,986 research outputs found

    Phylogenetic Modeling of the Histone Deacetylase 2-FK506 Binding (HD2-FKBP) Protein Family

    No full text
    Histone deacetylase 2-FKS06 binding proteins share two well-defined structural domains: the HD2 domain and the FKS06-binding domain. The evolutionary history of this protein family remains largely unknown. In this study, I examined phylogenetic relationships of members of the family using the Hidden Markov models CHMMs) that explain the similarities in the amino acid sequence. These HMMs were used to construct multiple sequence alignments and phylogenetic trees based on both the amino acid sequence of the entire proteins, and both the HD2 and the FKBP domains. The multiple sequence alignments and the phylogenetic trees show that similarities in phylogenetic relationship of the proteins in the family on two levels. These levels are the entire protein sequence and on level of both the domains, the HD2 domain and FKBP domain. These species are distributed through eukaryotes, being found in fungus, insects and plants. Phylogenetic analysis of these alignments suggests that both domains evolved together and all HD2-FKBP proteins evolved from a common ancestor. The goal of this project is to identify additional species that have proteins belonging to the HD2- FKBP family to show the evolutionary relationship between these proteins by comparing multiple sequence alignments and phylogenetic trees of the entire protein and each of the two domains. The domains were extracted using hidden Markov models of the domains. In order to extract the HD2 domain, a hidden Markov model of the domain was created. Using the hidden Markov models for the HD2 domain and the FKBP domain 39 new proteins from 24 new species were identified. In addition to fungus, insects, and higher plants which were previously classified as having proteins in this family, purple sea urchins and slime molds were added to the groups known to have proteins that belong to this family

    Using Software Transactional Memory in Interrupt-Driven Systems

    No full text
    Transactional memory presents a new concurrency control mechanism to handle synchronization between shared data. Dealing with concurrency issues has always been a difficulty when writing operating system software and using transactions aims to simplify matters. This thesis presents a framework for understanding how interrupt-driven device drivers can benefit from using transactional memory. A method for integrating software transactional memory (STM) into an operating system kernel is also developed and applied. This kernel uses STM over hardware transactional memory (HTM) because HTM requires modifications only implemented in simulated systems. By using STM, it is possible to build upon existing kernels and deploy operating systems onto commodity machines with communication peripherals. At the core is a modernized version of the Embedded Xinu operating system that has been ported to the Intel JA-32 architecture and modified to use a publicly available, production quality compiler and STM library available from Intel Corporation. The implementation of the Embedded Xinu kernel required several modifications to make use of the transactions offered by a library designed for use with user-level threads executing in Linux. Integrating the STM library into the kernel presents several challenges when dealing with a system that allows interrupts to enter at any time. By using transactions in device drivers, synchronization can be performed automatically and with greater granularity than traditional synchronizations methods. This may be used to to reduce the jitter-variations in interrupt handling-\u3c\u3eccurring in interrupt-driven device drivers when sharing data between the upper and lower halves, however this implementation does not show any conclusive results. This thesis discusses and presents the prototype implementation of Transactional Xinu. The prototype runs on standard hardware that exists today and provides a framework for future experimentation

    MeSH-Based Clustering of Biomedical Literature

    No full text
    The amount of online documents has grown tremendously in recent years that poses challenges for information retrieval from this vast collection. Text Mining, an application of machine learning addresses these challenges by providing techniques for information extraction from large text collections. One of the major areas of applications of text mining is biomedicine. The rapid growth of research in biomedical area is giving rise to a large number of literature published every year. It is difficult to keep pace with the current and related research in an area of interest. It is also difficult and time-consuming to read all the literature retrieved by a keyword search on a topic of interest. An efficient approach to address this problem is document clustering that generates meaningful groups of concepts which provide a better description of the data in a document collection. This study investigated document clustering of biomedical literature to identify concepts represented in large document collections. Biomedical literature is indexed by a controlled vocabulary, MeSH (Medical Subject Headings) which represent the major concepts discussed in a document. We compared the use of MeSH in representing the documents with that of full-text representation for document clustering

    Database Methods for Copy Number Variant Analysis of One Hundred Disease Associated Genes in Human Congenital Heart Disease

    Get PDF
    Human genetic variation occurs more commonly than was recognized after the completion of the Human Genome Sequencing Project in 2003. Submicroscopic human DNA analysis has revealed copy number variation (CNV) as the deletion or duplication of a genomic region potentially affecting gene dosage. Advanced genetic research now includes the study of CNVs in diseased subject groups compared to in house controls or online published datasets of control CNV data. Research labs choose from different bioinformatic algorithms to make the copy number calls. Solutions for further processing the copy number data into quantifiable form require collaboration with data analysts and include the use of relational databases. The aim of this thesis work was to develop a relational database solution for human copy number variation in subjects with cardiac malformations. The multipurpose database served as a central repository for the cohort demographic data as well as the entire experimental set of copy number variant data. Quantification and frequency analyses of the CNVs were executed via SQL queries. Database SQL queries generated raw data used for essential visualization tools including a detailed subject profile and a one hundred gene CNV spectra. The stated purpose of the study was to develop a descriptive analysis of genomic copy number associations in a well phenotyped congenital heart disease (CHD) population over one hundred disease associated genes. The relational database created to advance the research proved valuable in its data storage and retrieval capacity. Results showing consistency with published literature validated the accuracy of the query results generated for the CHD cohort

    Heuristics for Scaling up Distributed Protein Docking

    Get PDF
    Docking is a computational technique which predicts the interaction between a protein and a potential drug compound. Virtual screening is a tool, which employs docking, to investigate huge libraries of compounds and predicts potential drug molecules that bind favorably to the protein of interest. The size of one such commercially available library is about 13 million compounds. It would take approximately 400 years of CPU time to examine this library! As an alternative a high performance computing application with a distributed docking strategy is needed, which can efficiently predict the favorable compounds and can eventually be scaled for huge libraries. In this thesis, IncreDock, a scoring based incremental docking software has been developed to improve the efficiency of virtual screening process. IncreDock provides two approaches to the distributed docking problem. First, it allows for a completely parallel implementation where the entire library is explored simultaneously in an unordered fashion. Second, and more important, is the ordered incremental parallel implementation where the library is explored in increments and a scoring function is used to determine the order of dockings. IncreDock was used to perform docking studies on four different proteins using a library of 10,573 compounds. The results suggest that IncreDock is able to predict better compounds in early increments of dockings. IncreDock, thus, provides an effective initial strategy to sample out good ligands in less compute time and forms a good precursor to a tool that can proficiently investigate huge chemical libraries

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    A Temporal Difference Representation for Temporal Profiling of Short Time Series Microarray Gene Expression Data

    No full text
    A novel method for temporal profiling of short time series microarray data, inspired by reconstructed phase space (RPS) theory and the minimum entropy clustering algorithm is introduced. The augmented first difference space provides addition information, relating to the trend of the time series, for the clustering algorithm. This resultant higher dimension space is clustered using the minimum entropy clustering algorithm. The means of the subsequent clusters are then translated into a trend matrix, and a k-means clustering algorithm is applied to group the trends together. The results showed the proposed method was able to group genes with similar trends together and comparing with current approaches this new approach is able to find tighter clusters within the data

    Online Analytical Processing (OLAP) for Gene Expression Analysis

    No full text
    Each cell in a body contains thousands of genes encoded in DNA. At any given time, a cell expresses only a part of these genes as RNA transcripts. Gene expression is a measure of genes transcribed into RNA. Gene expression can be investigated in various experimental conditions such as by different tissues (e.g., normal vs diseased tissue), by developmental stages (e.g., early vs late development), drug responses (e.g., with vs without drug treatment), or disease states (e.g., breast cancer vs prostate cancer). Therefore, studies of gene expression address a variety of biological questions. Our research investigates an interactive tool for analyzing data from gene expression studies. Several technologies are used to measure gene expression. Among them, microarray technology is used broadly. A microarray, also called a gene chip or DNA chip, appears as a glass slide or nylon membrane on which genes are spotted in arrays. Each spot contains tens of millions of DNA molecules. During a microarray process, genes are labeled with fluorescence and scanned as an image. Through analyzing an image from the microarray technology, researchers can monitor the measurement of the expression levels of thousands of genes simultaneously. Microarray experiments provide large am,o unts of data for biological researchers. To obtain significant information from these large data sets, an appropriate method or tool is needed. There exist several methods or tools for microarray data analysis such as classification systems, clustering methods. In contrast with those methods, online analytical processing (OLAP) allows users to explore data interactively. Thus, we focus on building an OLAP tool for gene expression analysis. OLAP techniques are very useful since they provide excellent support for multidimensional views of the data and multiple hierarchies of dimensions. Moreover, they support three main operations: roll up, drill down and pivot. The combination of these three operations allows users to explore between summarized data and detailed one on any dimension and presents the results as a two-dimensional chart. We built a database by utilizing the hierarchies of gene products provided by the Gene Ontology (GO) consortium. The data sets used in this project were from the microarray experiments in breast cancer and prostate cancer categories of Standford Microarray Database (SMD). The project is developed on Mondrian OLAP server written in Java, with data stored in a local relational database (PostgreSQL) and a local Tomcat web server. The end system suggests that OLAP technology provide an interactive and fast way for gene expression analysis
    corecore