1,721,087 research outputs found

    New Algorithms for EST clustering

    No full text
    Philosophiae Doctor - PhDExpressed sequence tag database is a rich and fast growing source of data for gene expression analysis and drug discovery. Clustering of raw EST data is a necessary step for further analysis and one of the most challenging problems of modem computational biology. There are a few systems, designed for this purpose and a few more are currently under development. These systems are reviewed in the "Literature and software review". Different strategies of supervised and unsupervised clustering are discussed, as well as sequence comparison techniques, such as based on alignment or oligonucleotide compositions. Analysis of potential bottlenecks and estimation of computation complexity of EST clustering is done in Chapter 2. This chapter also states the goals for the research and justifies the need for new algorithm that has to be fast, but still sensitive to relatively short (40 bp) regions of local similarity. A new sequence comparison algorithm is developed and described in Chapter 3. This algorithm has a linear computation complexity and sufficient sensitivity to detect short regions of local similarity between nucleotide sequences. The algorithm utilizes an asymmetric approach, when one of the compared sequences is presented in a form of oligonucleotide table, while the second sequence is in standard, linear form. A short window is moved along the linear sequence and all overlapping oligonucleotides of a constant length in the frame are compared for the oligonucleotide table. The result of comparison of two sequences is a single figure, which can be compared to a threshold. For each measure of sequence similarity a probability of false positive and false negative can be estimated. The algorithm was set up and implemented to recognize matching ESTs with overlapping regions of 40bp with 95% identity, which is better than resolution ability of contemporary EST clustering tools This algorithm was used as a sequence comparison engine for two EST clustering programs, described in Chapter 4. These programs implement two different strategies: stringent and loose clustering. Both are tested on small, but realistic benchmark data sets and show the results, similar to one of the best existing clustering programs, 02_cluster, but with a significant advantage in speed and sensitivity to small overlapping regions of ESTs. On three different CPUs the new algorithm run at least two times faster, leaving less singletons and producing bigger clusters. With parallel optimization this algorithm is capable of clustering millions of ESTs on relatively inexpensive computers. The loose clustering variant is a highly portable application, relying on third-party software for cluster assembly. It was built to the same specifications as 02_ cluster and can be immediately included into the STACKPack package for EST clustering. The stringent clustering program produces already assembled clusters and can apprehend alternatively processed variants during the clustering process

    Ancient Genes in Cancer Gene Expression?

    No full text
    >Magister Scientiae - MScBacksround: The Cancer/testis (CT) antigens are a division of germ cell specific genes not expressed in somatic cells, exceptions being placental cells and 20Vo - 4OVo of cancer types. The aptitude of CT antigens to elicit humoral immune responses, their restricted expression profile, absence of major histocompatability complex expression in male germline cells have contributed to the emergent attraction of CT antigens as ideal, prospective cancer vaccination candidates. Motivation: Presently there are M CT gene families containing a total of 97 gene products and isoforms. Due to the promulgation in sensitivity and specificity of rapid serological immunodetection assays e.g. serial analysis of recombinant cDNA expression libraries (SEREX), the magnitude of novel CT genes and gene families will increase. Hence, characteization of this unique subset of CT genes is fundamental to our erudition of this rapidly emerging novel subset of genes. Obiectives: The sequencing of the human genome provides a useful biological framework for the categoization and systematization of rapidly accumulating biological information. A genomic approach was used to ascertain the locations of the CT genes in the human genome and determine if the genomic locations of the CT genes is nonrandom. An in-silico expression study was conducted for the CT genes with the aim of establishing if CT gene expression is restricted to the testis. A portion of the human genome housing the largest proportion of the CT genes was selected for analysis in order to determine if the surrounding genomic architecture influences CT gene expression. A comparative genomics approach was used in determining if the CT genes are "ancient genes

    "Development and implementation of ontology-based systems for mammalian gene expression profiling"

    No full text
    Philosophiae Doctor - PhDThe use of ontologies in the mapping of gene expression events provides an effective and comparable method to determine the expression profile of an entire genome across a large collection of experiments derived from different expression sources. In this dissertation I describe the development of the developmental human and mouse e VOC ontologies and demonstrate the ontologies by identifying genes showing a bias for developmental brain expression in human and mouse, identifying transcription factor complexes, and exploring the mouse orthologs of human cancer/testis genes. Model organisms represents fundamental aspects of mammal biology phenomena between model organism is complex and it is to be the meaningful, a simplified representation can be a powerful means for comparison illustrated here in two ways. Firstly, the ontologies have been used to illustrate methods to determine clusters of genes showing tissue-restricted expression in humans. The identification of tissue-restricted genes within an organism serves as an indication of the finetuning in the regulation of gene expression in a given tissue. Secondly, due to the differences in human and mouse gene expression on a temporal and spatial level, the ontologies were used to identify mouse orthologs of human cancer/testis genes showing cancer/testis characteristics. With the use of model systems such as mouse in the development of gene-targeted drugs in the treatment of disease, it is important to establish that the expression characteristics and profiles of a drug target in the model system is representative of the characteristics of the target in the system for which it is intended

    Development and implementation of ontology-based systems for mammalian gene expression profiling

    No full text
    Philosophiae Doctor - PhDThe use of ontologies in the mapping of gene expression events provides an effective and comparable method to determine the expression profile of an entire genome across a large collection of experiments derived from different expression sources. In this dissertation I describe the development of the developmental human and mouse eVOC ontologies and demonstrate the ontologies by identifying genes showing a bias for developmental brain expression in human and mouse, identifying transcription factor complexes, and exploring the mouse orthologs of human cancer/testis genes.Model organisms represent an important resource for understanding the fundamental aspects of mammalian biology. Mapping of biological phenomena between model organisms is complex and if it is to be meaningful, a simplified representation can be a powerful means for comparison. The implementation of the ontologies has been illustrated here in two ways.Firstly, the ontologies have been used to illustrate methods to determine clusters of genes showing tissue-restricted expression in humans. The identification of tissue restricted genes within an organism serves as an indication of the finetuning in the regulation of gene expression in a given tissue. Secondly, due to the differences in human and mouse gene expression on a temporal and spatial level, the ontologies were used to identify mouse orthologs of human cancer/testis genes showing cancer/testis characteristics. With the use of model systems such as mouse in the development of gene-targeted drugs in the treatment of disease, it is important to establish that the expression characteristics and profiles of a drug target in the model system is representative of the characteristics of the target in the system for which it is intended

    Development and implementation of ontology-based systems for mammalian gene expression profiling

    Get PDF
    Philosophiae Doctor - PhDThe use of ontologies in the mapping of gene expression events provides an effective and comparable method to determine the expression profile of an entire genome across a large collection of experiments derived from different expression sources. In this dissertation I describe the development of the developmental human and mouse eVOC ontologies and demonstrate the ontologies by identifying genes showing a bias for developmental brain expression in human and mouse, identifying transcription factor complexes, and exploring the mouse orthologs of human cancer/testis genes.Model organisms represent an important resource for understanding the fundamental aspects of mammalian biology. Mapping of biological phenomena between model organisms is complex and if it is to be meaningful, a simplified representation can be a powerful means for comparison. The implementation of the ontologies has been illustrated here in two ways.Firstly, the ontologies have been used to illustrate methods to determine clusters of genes showing tissue-restricted expression in humans. The identification of tissue restricted genes within an organism serves as an indication of the finetuning in the regulation of gene expression in a given tissue. Secondly, due to the differences in human and mouse gene expression on a temporal and spatial level, the ontologies were used to identify mouse orthologs of human cancer/testis genes showing cancer/testis characteristics. With the use of model systems such as mouse in the development of gene-targeted drugs in the treatment of disease, it is important to establish that the expression characteristics and profiles of a drug target in the model system is representative of the characteristics of the target in the system for which it is intended

    Development and implementation of ontology-based systems for mammalian gene expression profiling

    No full text
    >Magister Scientiae - MScThe use of ontologies in the mapping of gene expression events provides an effective and comparable method to determine the expression profile of an entire genome across a large collection of experiments derived from different expression sources. In this dissertation I describe the development of the developmental human and mouse e voe ontologies and demonstrate the ontologies by identifying genes showing a bias for developmental brain expression in human and mouse, identifying transcription factor complexes, and exploring the mouse orthologs of human cancer/testis genes

    Synergistic use of promoter prediction algorithms: A choice for small training dataset?

    Get PDF
    Philosophiae Doctor - PhDThis chapter outlines basic gene structure and how gene structure is related to promoter structure in both prokaryotes and eukaryotes and their transcription machinery. An in-depth discussion is given on variations types of the promoters among both prokaryotes and eukaryotes and as well as among three prokaryotic organisms namely, E.coli, B.subtilis and Mycobacteria with emphasis on Mituberculosis. The simplest definition that can be given for a promoter is: It is a segment of Deoxyribonucleic Acid (DNA) sequence located upstream of the 5' end of the gene where the RNA Polymerase enzyme binds prior to transcription (synthesis of RNA chain representative of one strand of the duplex DNA). However, promoters are more complex than defined above. For example, not all sequences upstream of genes can function as promoters even though they may have features similar to some known promoters (from section 1.2). Promoters are therefore specific sections of DNA sequences that are also recognized by specific proteins and therefore differ from other sections of DNA sequences that are transcribed or translated. The information for directing RNA polymerase to the promoter has to be in section of DNA sequence defining the promoter region. Transcription in prokaryotes is initiated when the enzyme RNA polymerase forms a complex with sigma factors at the promoter site. Before transcription, RNA polymerase must form a tight complex with the sigma/transcription factor(s) (figure 1.1). The 'tight complex' is then converted into an 'open complex' by melting of a short region of DNA within the sequence involved in the complex formation. The final step in transcription initiation involves joining of first two nucleotides in a phosphodiester linkage (nascent RNA) followed by the release of sigma/transcription factors. RNA polymerase then continues with the transcription by making a transition from initiation to elongation of the nascent transcript

    HIV Subtype C Diversity: Analysis of the Relationship of Sequence Diversity to Proposed Epitope Locations

    No full text
    >Magister Scientiae - MScSouthern Africa is facing one of the most serious HIV epidemics. This project contributes to the HIVNET, Network for Prevention Trials cohort for vaccine development. HIV's biology and rapid mutation rate have made vaccine design difficult. We examined HIV-l subtype C diversity and how it relates to CTL epitope location along viral gag sequences. We found a negative correlation between codon sites under positive selection and epitope regions; suggesting epitope regions are evolutionarily conserved. It is possible that detected due to the reference regions, yet fail to be viral population. To test if CTL clustering is an we calculated differences between the gag codons and the a weak negative correlation, suggesting epitopes in less conserved regions maybe evading detection. Locating conserved and optimal epitopes that can be recognized by CTLs is essential for the design of vaccine reagents

    A comparative genomics approach towards classifying immunity-related proteins in the tsetse fly

    No full text
    >Magister Scientiae - MScTsetse flies (Glossina spp) are vectors of African trypanosome (Trypanosoma spp) parasites, causative agents of Human African trypanosomiasis (sleeping sickness) and Nagana in livestock. Research suggests that tsetse fly immunity factors are key determinants in the success and failure of infection and the maturation process of parasites. An analysis of tsetse fly immunity factors is limited by the paucity of genomic data for Glossina spp. Nevertheless, completely sequenced and assembled genomes of Drosophila melanogaster, Anopheles gambiae and Aedes aegypti provide an opportunity to characterize protein families in species such as Glossina by using a comparative genomics approach. In this study we characterize thioester-containing proteins (TEPs), a sub-family of immunity-related proteins, in Glossina by leveraging the EST data for G.morsitans and the genomic resources of D. melanogaster, A. gambiae as well as A.aegypti.A total of 17 TEPs corresponding to Drosophila (four TEPs), Anopheles (eleven TEPs) and Aedes aegypti (two TEPs) were collected from published data supplemented with Genbank searches. In the absence of genome data for G. morsitans, 124 000 G.morsitans ESTs were clustered and assembled into 18 413 transcripts (contigs and singletons). Five Glossina contigs (Gmcn1115, Gmcn1116, Gmcn2398, Gmcn2281 and Gmcn4297) were identified as putative TEPs by BLAST searches. Phylogenetic analyses were conducted to determine the relationship of collected TEP proteins.Gmcn1115 clustered with DmtepI and DmtepII while Gmcn2398 is placed in a separate branch, suggesting that it is specific to G. morsitans.The TEPs are highly conserved within D. melanogaster as reflected in the conservation of the thioester domain, while only two and one TEPs in A. gambiae and A. aegypti thioester domain show conservation of the thioester domain suggesting that these proteins are subjected to high levels of selection. Despite the absence of a sequenced genome for G. morsitans, at least two putative TEPs where identified from EST data

    A comparative genomics approach towards classifying immunity-related proteins in the tsetse fly

    No full text
    >Magister Scientiae - MScTsetse flies (Glossina spp) are vectors of African trypanosome (Trypanosoma spp) parasites, causative agents of Human African trypanosomiasis (sleeping sickness) and Nagana in livestock. Research suggests that tsetse fly immunity factors are key determinants in the success and failure of infection and the maturation process of parasites. An analysis of tsetse fly immunity factors is limited by the paucity of genomic data for Glossina spp. Nevertheless, completely sequenced and assembled genomes Drosophila melanogaster, Anopheles gambiae and Aedes aegypti provide an opportunity to characterize protein families in species such as G/ossiza by using a comparative genomics approach. In this study, we characterize thioester-containing proteins (TEPs), a sub-family of immunity-related proteins, in Glossinaby leveraging the EST data for G. morsitans and the genomic resources of D. melanogaster, A. gambiae as well as A. aegypt
    corecore