1,721,171 research outputs found
Methods for efficient storage and indexing in XML databases.
As the eXtensible Markup Language (XML) continues to increase in popularity, it is clear that large repositories of XML data sets will emerge in the near future. Existing techniques for XML query processing are not very efficient and unlikely to scale well with large data sets. In this thesis, we address these shortcomings and investigate various aspects of query processing on large XML data sets. In the first part of the thesis, we propose an algorithm, called XORator, that maps documents based on their schema information into constructs in an Object-Relational Database Management System (ORDBMS). We compare the effectiveness of the XORator algorithm with an algorithm that maps XML data in a Relational Database Management System (RDBMS) and show that the XORator technique results in significant improvements in query response times. In the second part of this thesis, we propose a technique, called PAID, for storing XML data, independent of XML schemas. The PAID technique extends a previous numbering scheme by including path information and a pointer to the node's parent element. Our experimental results demonstrate that the PAID technique is more efficient than existing strategies by several orders of magnitude. The third part of this thesis presents an XML Index Selection Tool (XIST) that examines a combination of database query workload, data statistics, and XML schemas to suggest a set of indices that are beneficial to build. Our experiments show that XIST produces index recommendations that are more efficient than those produced by existing techniques. The final part of this thesis presents a micro-benchmark, called the Michigan benchmark, that can be used to evaluate the performance of XML database systems. The benchmark is an engineers' benchmark and is designed to pinpoint the strengths and weaknesses of individual components that constitute the entire DBMS. We have used the benchmark to test three databases and understand the factors that are critical to the performance in these DBMSs. Collectively, our research presents techniques for efficiently managing large XML data repositories. Although the bulk of the thesis has focused on techniques applicable to commercial ORDBMSs, many of these methods, such as the XIST tool, can also be adapted for native XML systems.PhDApplied SciencesComputer scienceElectrical engineeringUniversity of Michigan, Horace H. Rackham School of Graduate Studieshttp://deepblue.lib.umich.edu/bitstream/2027.42/123943/2/3106155.pd
Intelligent Ship Arrangements: A New Approach to General Arrangement
A new surface ship general arrangement optimization system developed at the University of Michigan is described. The Intelligent Ship Arrangements system is a native C++, Leading Edge Architecture for Prototyping Systems-compatible software system that will assist the designer in developing rationally based arrangements that satisfy design specific needs as well as general Navy requirements and standard practices to the maximum extent practicable. This software system is intended to be used following or as a latter part of ASSET synthesis. The arrangement process is approached as two essentially two-dimensional tasks. First, the spaces are allocated to Zone-decks, one deck in one vertical zone, on the ship's inboard profile. Then the assigned spaces are arranged in detail on the deck plan of each Zone-deck in succession. Consideration is given to overall location, adjacency, separation, access, area requirements, area utilization, and compartment shape. The system architecture is quite general to facilitate its evolution to address additional design issues, such as distributive system design, in the future
Architecture-conscious storage management.
Designers of database management systems (DBMS) have traditionally focussed on alleviating the disk I/O performance bottleneck. Because of the decreasing price and the increasing capacity of random access memory, systems with large main-memory configurations are becoming more prevalent. As the amount of main memory grows larger, data sets reside in main memory longer, and the disk I/O bottleneck shifts to the main-memory hierarchy. Because of this shift, DBMSs must become aware of the underlying architecture to maximize performance. This dissertation explores architecture-conscious design techniques that exploit the underlying architecture to improve DBMS performance. The first contribution of this thesis is an architecture-conscious data-storage technique, called Data Morphing (DM). The DM process dynamically analyzes the query workload and reorganizes the data to improve its spatial locality. The improved spatial locality reduces the number of processor cache misses that occur during query processing, thereby significantly improving performance. The second contribution of this thesis is an analysis of two architecture-conscious secondary index structures: the CSB+-tree and an extendible hash index. Analytical models based on the underlying microarchitectural behavior are developed for both indexes. Analysis of the CSB+-tree index shows that using a larger node size than originally proposed improves performance. Similar analysis of the extendible hash index shows that the overflow chain length is more critical than bucket size. Because of the tight coupling between DBMS performance and the underlying architecture, research for future processor design must take place in the simulation domain. Unfortunately, simulating large-scale database workloads is difficult due to their high complexity and cost. The third contribution of this thesis shows that the architectural behavior of a large scale database workload can be approximated by a much smaller workload. Using this smaller workload is much more conducive to simulation because of the reduced complexity and cost. This research is on the cusp of a new effort in main-memory database research. The techniques presented within are not only applicable on current DMBSs, but apply to future systems as well.PhDApplied SciencesComputer scienceUniversity of Michigan, Horace H. Rackham School of Graduate Studieshttp://deepblue.lib.umich.edu/bitstream/2027.42/124420/2/3138166.pd
Towards a unified framework for efficient access methods and query operations in spatio-temporal databases.
Spatio-temporal databases are required to efficiently support queries on large numbers of continuously moving objects. In this thesis we propose three techniques to address this challenge. We develop the STRIPES indexing method, which indexes predicted trajectories in a dual transformed space. Trajectories for objects in d-dimensional space are transformed into points in 2d-dimensional space and are indexed with a hierarchical grid decomposition structure. STRIPES can evaluate a range of queries including time-slice, window, and moving queries. Extensive experimental evaluation shows that STRIPES is significantly faster than leading existing predicted trajectory index (the TPR*-tree) for both updates and queries. The All Nearest Neighbor (ANN) operation is a commonly used primitive for analyzing large multi-dimensional datasets. Traditional R*-tree based methods use a pruning metric called MAXMAXDIST, which allows the algorithms to prune nodes in the index that need not be examined during the ANN computation. We introduce a new pruning metric called NXNDIST, and show that this metric is far more effective than MAXMAXDIST. We also propose the MBRQT index structure, and using extensive experimental evaluation show that MBRQT offers better speedup in ANN computation than the commonly used R*-tree index. In addition, we present the MBA algorithm using depth-first index traversal and bi-directional node expansion. Furthermore, we extend our method to evaluate the more general All-k-Nearest-Neighbor (AkNN) operation. Traditional spatial and traditional temporal joins have been widely studied in the past, but there is very little work on the more complex problem of trajectory joins. We present a general framework called JiST, that introduces a broad class of trajectory join operations and offers a set of algorithms to efficiently evaluate these operations. We then introduce the notion of Trajectory Privacy, and show the application of the JiST framework in the context of privacy preservation. Finally, we present detailed experimental results that demonstrate the efficiency and scalability of the JiST join algorithms. To the best of our knowledge, JiST is the first comprehensive framework for complex trajectory join operations and paves the foundation for building a complex querying platform for emerging trajectory based applications.PhDApplied SciencesComputer scienceUniversity of Michigan, Horace H. Rackham School of Graduate Studieshttp://deepblue.lib.umich.edu/bitstream/2027.42/127006/2/3304944.pd
Diversity oriented design of various hydrazides and their in vitro evaluation against Mycobacterium tuberculosis H37Rv strains
Control and prevention of tuberculosis is a major challenge, as one-third of the world’s population is infected with Mycobacterium tuberculosis. The resurgence of tuberculosis and the emergence of multidrug-resistance strains of mycobacteria, necessitate the search for new class of antimycobacterial agents. As a part of investigation of new antitubercular agents in this laboratory, we describe the syntheses of various hydrazides of comarins, quinolones and pyrroles and screening against M. tuberculosis (Mtb) H37Rv by using rifampin as a standard drug. Among the designed molecules, the most prominent compounds 2a–g, 4a and 9a showed >90% GI at MIC <6.25 μg/mL. Finally, these studies suggests that compounds 2a–g, 4a and 9a may serve as promising lead scaffolds for further generation of new anti-TB agents
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
