1,721,005 research outputs found
Analysis and application of hash-based similarity estimation techniques for biological sequence analysis
In Bioinformatics, a large group of problems requires the computation or estimation of sequence similarity. However, the analysis of biological sequence data has, among many others, three capital challenges: a large amount generated data which contains technology-specific errors (that can be mistaken for biological signals), and that might need to be analyzed without access to a reference genome. Through the use of locality sensitive hashing methods, both the efficient estimation of sequence similarity and tolerance against the errors specific to biological data can be achieved.
We developed a variant of the winnowing algorithm for local minimizer computation, which is specifically geared to deal with repetitive regions within biological sequences. Through compressing redundant information, we can both reduce the size of the hash tables required to save minimizer sketches, as well as reduce the amount of redundant low quality alignment candidates.
Analyzing the distribution of segment lengths generated by this approach, we can better judge the size of required data structures, as well as identify hash functions feasible for this technique.
Our evaluation could verify that simple and fast hash functions, even when using small hash value spaces (hash functions with small codomain), are sufficient to compute compressed minimizers and perform comparable to uniformly randomly chosen hash values. We also outlined an index for a taxonomic protein database using multiple compressed winnowings to identify alignment candidates. To store MinHash values, we present a cache-optimized implementation of a hash table using Hopscotch hashing to resolve collisions.
As a biological application of similarity based analysis, we describe the analysis of double digest restriction site associated DNA sequencing (ddRADseq). We implemented a simulation software able to model the biological and technological influences of this technology to allow better development and testing of ddRADseq analysis software. Using datasets generated by our software, as well as data obtained from population genetic experiments, we developed an analysis workflow for ddRADseq data, based on the Stacks software. Since the quality of results generated by Stacks strongly depends on how well the used parameters are adapted to the specific dataset, we developed a Snakemake workflow that automates preprocessing tasks while also allowing the automatic exploration of different parameter sets. As part of this workflow, we developed a PCR deduplication approach able to generate consensus reads incorporating the base quality values (as reported by the sequencing device), without performing an alignment first.
As an outlook, we outline a MinHashing approach that can be used for a faster and more robust clustering, while addressing incomplete digestion and null alleles, two effects specific for ddRADseq that current analysis tools cannot reliably detect
Analysis and application of evolutionary processes to tackle HIV-1 entry
Im Laufe der Jahrmillionen hat die Evolution durch einige einfache Mechanismen wie Mutation, Selektion oder auch Vererbung eine erstaunliche Artenvielfalt hervorgebracht. Diese Prinzipien können auch beim computergestützten Entwurf von Proteinen und/oder Proteinsequenzen mit gewünschten Eigenschaften, wie z.B. Stabilität
oder Funktionalität einer Proteinstruktur, angewandt werden. Da jedoch der mögliche Konformations- und Sequenzraum für bereits kleine Proteine immens groß wird, werden hier vereinfachte Gitterproteinmodelle verwendet. Im ersten Teil der Promotionsarbeit werden evolutionäre Algorithmen, im Besonderen S Metric Selection - Evolutionary Multi-objective Optimisation Algorithm (SMS-EMOA), implementiert und angewandt um möglichst optimale evolutionäre Parameter zu identifizieren, z.B. Populationsgröße oder Mutationsrate. Interessanterweise spielt die richtige Auswahl der evolutionären Parameter eine entscheidende Rolle bezüglich der Effizienz der Algorithmen. Im zweiten Teil der Arbeit wird die Evolution von Proteinen beobachtet und analysiert. Ein besonderes Augenmerk wird dabei auf Positionen gelegt, die nicht konserviert sind. Gleichwohl können diese mit kompensatorischen Mutationen an anderen Stellen im Protein strukturell wichtige Funktionen einnehmen. Hierbei werden verschiedene Koevolutionsmethoden, wie z.B. die Mutual Information (MI) oder die Direct Coupling Analysis (DCA), weiterentwickelt und verglichen.
Anschließend wird die DCA-Methode mit einer neu verbesserten Gewichtung angewandt um koevolvierende Positionen im Humanen Immundefizienz-Virus (HIV) Hüllprotein-Komplex (Env) vorherzusagen. Bemerkenswerterweise wurden dabei sowohl bereits in der Literatur beschriebene als auch noch unbekannte Positionen identifiziert, die eine entscheidende Rolle im Eintritt des Viruses in die humane Wirtszelle spielen können. Schließlich wurden die koevolvierenden Positionen bei der Erstellung eines Homologiemodells des Protein-Komplexes verwendet
Parallelization, scalability, and reproducibility in next generation sequencing analysis
The analysis of next-generation sequencing (NGS) data is a major topic in bioinformatics: short reads obtained from DNA, the molecule encoding the genome of living organisms, are processed to provide insight into biological or medical questions. This thesis provides novel solutions to major topics within the analysis of NGS data, focusing on parallelization, scalability and reproducibility. The read mapping problem is to find the origin of the short reads within a given reference genome. We contribute the q-group index, a novel data structure for read mapping with particularly small memory footprint. The q-group index comes with massively parallel build and query algorithms targeted towards modern graphics processing units (GPUs). On top, the read mapping software PEANUT is presented, which outperforms state of the art read mappers in speed while maintaining their accuracy. The variant calling problem is to infer (i.e., call) genetic variants of individuals compared to a reference genome using mapped reads. It is usually solved in a Bayesian way. Often, variant calling is followed by filtering variants of different biological samples against each other. With state of the art solutions, the filtering is decoupled from the calling, leading to difficulties in controlling the false discovery rate. In this work, we show how to integrate the filtering into the calling with an algebraic approach and provide an intuitive solution for controlling the false discovery rate along with solving other challenges of variant calling like scaling with a growing set of biological samples. For this, a hierarchical index data structure for storage of preprocessing results is presented and compression strategies are provided. The developed methods are implemented in the software ALPACA. Depending on the research question, the analysis of NGS data entails many other steps, typically involving diverse tools, data transformations and aggregation of results. These steps can be orchestrated by work ow management. We present the general purpose work ow system Snakemake, which provides an easy to read domain-specific language for defining and documenting work ows, thereby ensuring reproducibility of analyses. The language is complemented by an execution environment that allows to scale a work ow to available resources, including parallelization across CPU cores or cluster nodes, restricting memory usage or the number of available coprocessors like GPUs. The benefits of using Snakemake are exemplified by combining the presented approaches for read mapping and variant calling to a complete, scalable and reproducible NGS analysis
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Chemometrics and statistical analysis in raman spectroscopy-based biological investigations
As mentioned in the chapter 1, chemometrics has become an essential tool in Raman spectroscopy-based biological investigations and significantly enhanced the sensitivity of Raman spectroscopy-based detection. However, there are some open issues on applying chemometrics in Raman spectroscopy-based biological investigations. An automatic proce- dure is needed to optimize the parameters of the mathematical baseline correction. Spectral reconstruction algorithm is required to recover a fluorescence-free Raman spectrum from the two Raman spectra measured with different excitation wavelengths for the shifted-excitation Raman difference spectroscopy (SERDS) technique. Guidelines are necessary for reliable model optimization and rigorous model evaluation to ensure high accuracy and robustness in Raman spectroscopy-based biological detection. Computational methods are required to enable a trained model to successfully predict new data that is significantly different from the training data due to inter-replicate variations. These tasks were tackled in this thesis. The related investigations were related to three main topics: baseline correction, statistical modeling, and model transfer.Wie im Kapitel 1 erwähnt, ist die Chemometrie zu einem essentiellen Werkzeug für biolo- gische Untersuchungen mittels der Raman-Spektroskopie geworden und hat die Sensitivität der Raman-spektroskopischen Detektion erheblich verbessert. Es gibt jedoch einige offene Fragen, welche die Anwendung der Chemometrie in Raman-spektroskopischen Untersuchun- gen biologischer Proben betreffen. Zum Beispiel wird eine automatische Prozedur benötigt, um die Parameter einer mathematischen Basislinienkorrektur zu optimieren. Ein SERDS- Rekonstruktionsalgorithmus ist erforderlich, um ein Fluoreszenz-freies Raman-Spektrum aus den zwei Raman-Spektren zu extrahieren, welche bei der Shifted-excitation-Raman-Differenz- Spektroskopie (SERDS) gemessen werden. Des Weiteren sind Richtlinien erforderlich, welche eine zuverlässige Modelloptimierung und eine rigorose Modellevaluation erlauben. Durch diese Richtlinien wird eine hohe Genauigkeit und Robustheit der Raman-spektroskopischen Detektion biologischer Proben gewährleistet. Computergestützte Methoden sind nötig, um mit einem trainierten Modell erfolgreich neue Daten, die sich aufgrund von Inter-Replikat- Variationen signifikant von den Trainingsdaten unterscheiden, vorherzusagen. Diese vier Probleme sind Beispiele für offene Fragen in der Chemometrie und diese vier Probleme wur- den in dieser Arbeit behandelt. Die damit verbundenen Untersuchungen bezogen sich auf drei Hauptthemen: die Basislinienkorrektur, die statistische Modellierung und der Modell- transfer
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
