Computing and Informatics (E-Journal - Institute of Informatics, SAS, Bratislava)
Not a member yet
1506 research outputs found
Sort by
Managing Uncertain Mediated Schema and Semantic Mappings Automatically in Dataspace Support Platforms
Contrary to existing heterogeneous data integration systems which need to be fully integrated before using, a Dataspace Support Platform is a self-sustained system which automatically provides for the user its best endeavor results regardless of how integrated its sources are. Therefore, a Dataspace Support Platform needs to support uncertainty in mediated schema and in schema mappings. This paper proposes a novel approach to automatically providing reliable mediated schemas and reliable semantic mappings in Dataspace Support Platforms. Our aim is to increase the system's endeavor results by leading it to considering as much as possible information available in any source connected. In fact, we first extract from the source schemas, their corresponding graph representations. Then, we introduce algorithms which automatically extract a set of mediated schemas from the graph representations and a set of semantic mappings between a source and a target mediated schema. Finally, we assign reliability degrees to the mediated schema generated and to the semantic mappings. Indeed, the higher the reliability degree of a given mediated schema or semantic mapping, the more consistent with the source it is. Compared with existing systems, experimental results show that our system is faster and, although completely automatic, it produces reliable mediated schemas and reliable semantic mappings which are as accurate as those produced by semi-automatic systems
Approaches to samples selection for machine learning based classification of textual data
The paper focuses on the process of selecting representative sample documents written in a natural language that can be used as the basis for automatic selection or classification of textual documents. A method of selecting the examples from a larger set of candidate examples, called automatic biased sample selection, is compared to random and manual selection. The methods are evaluated by experiments carried out with real world data consisting of customer reviews, with different document representations and similarity measures. Presented approach, that provided satisfactory results, faces problems related to processing user created content and huge computational complexity and can be used as an alternative to manual selection and evaluation of textual samples
An Efficient Method of Summarizing Documents Using Impression Measurements
Automatic generic document summarization based on unsupervised schemes is a very useful approach because it does not require training data. Although techniques using latent semantic analysis (LSA) and non-negative matrix factorization (NMF) have been applied to determine topics of documents, there are no researches on reduction of matrix and speeding up of computation of the NMF method. In order to achieve this scheme, this paper utilizes the generic impressive expressions from newspapers to extract important sentences as summary. Therefore, it has no stemming processes and no filtering of stop words. Generally, novels are typical documents providing sentimental impression for readers. However, newspapers deliver different impressions for new knowledge because they inform readers about current events, informative articles and diverse features. The proposed method introduces impressive expressions for newspapers and their measurements are applied to the NMF method. From 100 KB text data of experimental results by the proposed method, it turns out that the matrix size reduces by 80 % and the computation of the NMF method becomes 7 times faster than with the original method, without degrading the relevancy of extracted sentences
Multiple Route Generation Using Simulated Niche Based Particle Swarm Optimization
This research presents an optimization technique for multiple routes generation using simulated niche based particle swarm optimization for dynamic online route planning, optimization of the routes and proved to be an effective technique. It effectively deals with route planning in dynamic and unknown environments cluttered with obstacles and objects. A simulated niche based particle swarm optimization (SN-PSO) is proposed using modified particle swarm optimization algorithm for dealing with online route planning and is tested for randomly generated environments, obstacle ratio, grid sizes, and complex environments. The conventional techniques perform well in simple and less cluttered environments while their performance degrades with large and complex environments. The SN-PSO generates and optimizes multiple routes in complex and large environments with constraints. The traditional route optimization techniques focus on good solutions only and do not exploit the solution space completely. The SN-PSO is proved to be an efficient technique for providing safe, short, and feasible routes under dynamic constraints. The efficiency of the SN-PSO is tested in a mine field simulation with different environment configurations and successfully generates multiple feasible routes
Data Integration in Mediated Service Compositions
A major aim of the Web service platform is the integration of existing software and information systems. Data integration is a central aspect in this context. Traditional techniques for information and data transformation are, however, not sufficient to provide flexible and automatable data integration solutions for Web and Cloud service-enabled information systems. The difficulties arise from a high degree of complexity in data structures in many applications and from the additional problem of heterogeneity of data representation in applications that often cross organisational boundaries. We present an integration technique that embeds a declarative data transformation technique based on semantic data models as a mediator service into a Web service-oriented information system architecture. Automation through consistency-oriented semantic data models and flexibility through modular declarative data transformations are the key enablers of the approach. Automation is needed to enable dynamic integration and composition. Modifiability is another aim here that benefits from consistency and modularity
An Information- Theoretical Model for Streaming Media Based Stegosystems
Steganography in streaming media differs from steganography in images or audio files because of the continuous embedding process and the necessary synchronization of sender and receiver due to packet loss in streaming media. The conventional theoretical model for image steganography is not appropriate for explaining the security scenarios for streaming media based stegosystems. In this paper, we propose a new information-theoretical model with two pseudo-random sequences imitating the continuous embedding and synchronization characteristics of streaming media based stegosystems. We also discuss the statistical properties of Voice over Internet Protocol (VoIP) speech streams through theoretical analysis and experimental testing. The experimental results show the bit stream consisting of fixed codebook parameters in speech frames is similar in statistical characteristics to a white-noise sequence. The relative entropy between the VoIP speech stream and the embedded secret message has been found to be zero. This leads us to conclude that the proposed streaming media based stegosystem is secure against statistical detection; in other words, the statistical measures cannot detect the existence of the secret message embedded in VoIP speech streams
Document Clustering with Bursty Information
Nowadays, almost all text corpora, such as blogs, emails and RSS feeds, are a collection of text streams. The traditional vector space model (VSM), or bag-of-words representation, cannot capture the temporal aspect of these text streams. So far, only a few bursty features have been proposed to create text representations with temporal modeling for the text streams. We propose bursty feature representations that perform better than VSM on various text mining tasks, such as document retrieval, topic modeling and text categorization. For text clustering, we propose a novel framework to generate bursty distance measure. We evaluated it on UPGMA, Star and K-Medoids clustering algorithms. The bursty distance measure did not only perform equally well on various text collections, but it was also able to cluster the news articles related to specific events much better than other models
An Intelligent Genetic Algorithm for Mining Classification Rules in Large Datasets
Genetic algorithm is a popular classification algorithm which creates a random population of candidate solutions and makes them to evolve into a suitable accurate solution for a given problem by processing them iteratively for several generations. During each generation the training data set is accessed by the genetic algorithm only for the population member's fitness calculation and no other extra knowledge about the problem domain is extracted from the training data set. Even the domain knowledge stored in the chromosome code of the population may be lost in the future generations due to genetic operations. All the genetic operations like crossover and mutation are probability based and they do not depend upon the domain knowledge. This phenomenon makes the genetic algorithm to converge slowly. This paper proposes a genetic algorithm which tries to gain maximum knowledge in between the generations and store them in the form of knowledge chromosomes. The gained knowledge is used to make predictions about the search space and to guide the search process to an area with potential solutions in the subsequent generations. This makes the genetic algorithm to converge quickly which in turn reduces the learning cost. The experiments show that the run time is reduced considerably when compared with the state-of-the-art evolutionary algorithm
Efficient Decimation of Polygonal Models Using Normal Field Deviation
A simple and robust greedy algorithm has been proposed for efficient and quality decimation of polygonal models. The performance of a simplification algorithm depends on how the local geometric deviation caused by a local decimation operation is measured. As normal field of a surface plays key role in its visual appearance, exploiting the local normal field deviation in a novel way, a new measure of geometric fidelity has been introduced. This measure has the potential to identify and preserve the salient features of a surface model automatically. The resulting algorithm is simple to implement, produces approximations of better quality and is efficient in running time. Subjective and objective comparisons validate the assertion. It is suitable for applications where the focus is better speed-quality trade-off, and simplification is used as a processing step in other algorithms
An Approach to Generating Arguments over DL-Lite Ontologies
Argumentation frameworks for ontology reasoning and management have received extensive interests in the field of artificial intelligence in recent years. As one of the most popular argumentation frameworks, Besnard and Hunter's framework is built on arguments in form of where Phi is consistent and minimal for entailing phi. However, the problem about generating arguments over ontologies is still open. This paper presents an approach to generating arguments over DL-Lite ontologies by searching support paths in focal graphs. Moreover, theoretical results and examples are provided to ensure the correctness of this approach. Finally, we show this approach has the same complexity as propositional revision