1,720,979 research outputs found
Building blocks for semantic data organization on the desktop
Die Organisation von (Multimedia-) Daten auf Desktop-Systemen wird derzeit
hauptsächlich durch das Einordnen von Dateien in ein hierarchisches Dateisystem
bewerkstelligt. Zusätzlich werden gewisse Inhalte (z.B. Musik oder Fotos) von
spezialisierter Software mit Hilfe Datei-bezogener Metadaten verwaltet. Diese
Metadaten werden meist direkt im Dateikopf in einer Unzahl verschiedener,
vorwiegend proprietärer Formate gespeichert. Allgemein nehmen Metadaten und
Links die Schlüsselrollen in fortgeschrittenen Datenorganisationskonzepten ein,
ihre eingeschränkte Unterstützung in vorherrschenden Dateisystemen macht die
Einführung solcher Konzepte auf dem Desktop jedoch schwierig: Erstens müssen
Anwendungen sowohl Dateiformat als auch Metadatenschema verstehen um auf
Metadaten zugreifen zu können; zweitens ist ein getrennter Zugriff auf Daten und
Metadaten nicht möglich und drittens kann man solche Metadaten nicht mit
mehreren Dateien oder mit Dateiordnern assoziieren obgleich letztere die derzeit
wichtigsten Konstrukte für die Dateiorganisation darstellen. Dies bedeutet in
weiterer Folge: (i) eingeschränkte Möglichkeiten der Datenorganisation, (ii)
eingeschränkte Navigationsmöglichkeiten, (iii) schlechte Auffindbarkeit der
gespeicherten Daten, und (iv) Fragmentierung von Metadaten. Obschon es Versuche
gab, diese Situation (zum Beispiel mit Hilfe semantischer Dateisysteme) zu
verbessern, wurden die meisten dieser Probleme bisher vor allem im Web und im
Speziellen im semantischen Web adressiert und gelöst. Das Anwenden dort
entwickelter Lösungen auf dem Desktop, einer zentralen Plattform der Daten- und
Metadatenmanipulation, wäre zweifellos von Vorteil.
In der vorliegenden Arbeit wird ein neues, rückwärts-kompatibles Metadatenmodell
als Lösungsversuch für die oben genannten Probleme präsentiert. Dieses Modell
basiert auf stabilen Datei-Identifikatoren und externen, semantischen, Datei-
bezogenen Metadatenbeschreibungen welche im RDF Graphenmodell repräsentiert
werden. Diese Beschreibungen sind durch eine einheitliche Linked-Data-
Schnittstelle zugänglich und können mit anderen Beschreibungen und Ressourcen
verlinkt werden. Im Speziellen erlaubt dieses Modell semantische Links zwischen
lokalen Dateisystemobjekten und Netzressourcen im Web sowie im entstehenden
“Daten Web” und ermöglicht somit die Integration dieser Datenräume. Das Modell
hängt entscheidend von der Stabilität dieser Links ab weshalb zwei Algorithmen
präsentiert werden, welche deren Integrität in lokalen und vernetzten Umgebungen
erhalten können. Dies bedeutet, dass Links zwischen Dateisystemobjekten,
Metadatenbeschreibungen und Netzressourcen nicht brechen wenn sich deren
Adressen ändern, z.B. wenn Dateien verschoben oder Linked-Data Ressourcen unter
geänderten URIs publiziert werden. Schließlich wird eine prototypische
Implementierung des vorgeschlagenen Metadatenmodells präsentiert, welche
demonstriert wie die Summe dieser Bausteine eine Metadatenschicht bildet die als
Grundlage für semantische Datenorganisation auf dem Desktop verwendet werden
kann.The organization of (multimedia) data on current desktop systems is done to a
large part by arranging files in hierarchical file systems, but also by
specialized applications (e.g., music or photo organizing software) that make
use of file-related metadata for this task. These metadata are predominantly
stored in embedded file headers, using a magnitude of mainly proprietary
formats. Generally, metadata and links play the key roles in advanced data
organization concepts. Their limited support in prevalent file system
implementations, however, hinders the adoption of such concepts on the desktop:
First, non-uniform access interfaces require metadata consuming applications to
understand both a file’s format and its metadata scheme; second, separate
data/metadata access is not possible, and third, metadata cannot be attached to
multiple files or to file folders although the latter are the primary constructs
for file organization. As a consequence of this, current desktops suffer, inter
alia, from (i) limited data organization possibilities, (ii) limited
navigability, (iii) limited data findability, and (iv) metadata fragmentation.
Although there were attempts to improve this situation, e.g., by introducing
semantic file systems, most of these issues were successfully addressed and
solved in the Web and in particular in the Semantic Web and reusing these
solutions on the desktop, a central hub of data and metadata manipulation, is
clearly desirable.
In this thesis a novel, backwards-compatible metadata model that addresses the
above-mentioned issues is introduced. This model is based on stable file
identifiers and external, file-related, semantic metadata descriptions that are
represented using the generic RDF graph model. Descriptions are accessible via a
uniform Linked Data interface and can be linked with other descriptions and
resources. In particular, this model enables semantic linking between local file
system objects and remote resources on the Web or the emerging Web of Data,
thereby enabling the integration of these data spaces. As the model crucially
relies on the stability of these links, we contribute two algorithms that
preserve their integrity in local and in remote environments. This means that
links between file system objects, metadata descriptions and remote resources do
not break even if their addresses change, e.g., when files are moved or Linked
Data resources are re-published using different URIs. Finally, we contribute a
prototypical implementation of the proposed metadata model that demonstrates how
these building blocks sum up to constitute a metadata layer that may act as a
foundation for semantic data organization on the desktop
A novel compression approach for mapped high-throughput sequencing data set
Eine der größten aktuellen Herausforderungen im Zusammenhang mit Hochdurchsatz-Sequenzierungsexperimenten (High-Throughput Sequencing, HTS) liegt nicht im Erzeugen der Daten selbst, sondern in deren Prozessierung, Speicherung und Übertragung. Die enorme Größe dieser Daten motiviert die Entwicklung
von Datenkompressionsalgorithmen für die Realisierung der verschiedenen Datenspeicherkonzepte die auf die produzierten (Zwischen-)Ergebnisse von HTS Experimenten angewandt werden.
Die vorliegende Arbeit gibt einen Überblick über das Feld der Hochdurchsatz-Nukleinsäure-Sequenzierung und in aktuelle Ansätze für die Kompression solcher Daten. Im Hauptteil der Arbeit wird NGC vorgestellt, ein Werkzeug für die Kompression von gemappten reads die im weitverbreiteten SAM Format
gespeichert sind (eine Art von HTS Daten). NGC ermöglicht sowohl verlustfreie als auch verlustbehaftete Kompression und beinhaltet zwei neuartige Ideen: Erstens enthält es eine Methode zur Reduktion der erforderlichen
Code-Wörter, welche gemeinsame Merkmale der reads die an dieselbe genomische Position gemappt wurden ausnützt. Zweitens beinhaltet NGC eine konfigurierbare Methode für die Quantisierung der Qualitätswerte welche deren Einfluss auf nach-gelagerte Anwendungen berücksichtigt.
NGC, mit mehreren echten Datensätzen evaluiert, spart 33-66% des benötigten Speicherplatzes bei verlustfreier und bis zu 98% des benötigten Speicherplatzes bei verlustbehafteter Kompression ein. Durch die Anwendung zweier gängiger Varianten- und Genotyp-Vorhersagewerkzeuge auf die dekomprimierten Daten
wird gezeigt, dass die verlustbehaftete Kompression, besser als vergleichbare Werkzeuge in manchen Konfigurationen, über 99% der gefundenen Varianten präserviert.A major challenge of current high-throughput sequencing (HTS) experiments is not only the generation of the sequencing data itself but also their processing, storage and transmission. The enormous size of these data motivates the development of data compression algorithms usable for the implementation of the various storage policies that are applied to the produced intermediate and final result files.
This thesis gives a brief introduction into the field of high-throughput nucleic acid sequencing and into current approaches for the compression of the data resulting from such experiments. In the main part of the thesis, NGC, a tool for the compression of mapped read data stored in the SAM format (one kind of HTS data), is presented. NGC enables lossless and lossy compression and introduces two novel ideas: First, it contains a way to reduce the number of required code words by exploiting common features of the sequenced reads mapped to the same genomic positions; second, it contains a highly configurable way for the quantization of per-base quality values which takes their influence on downstream analyses into account.
NGC, evaluated with several real-world data sets, saves 33-66% of disc space using lossless and up to 98% disc space using lossy compression. By applying two popular variant and genotype prediction tools to the decompressed data, we show that the lossy compression modes preserve over 99% of all called variants while outperforming comparable methods in some configurations
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
Author-wise bibliometric analysis based on entropy.
Author-wise bibliometric analysis based on entropy.</p
- …
