1,721,390 research outputs found
Landscape Analysis of Ontologies in Materials Science and Engineering
<p><span>Ontologies are widely used in the domain of materials science for the purpose of describing experiments, processes, properties of materials, and experimental and computational workflows. There are several online platforms for accessing and sharing ontologies related to materials science and engineering available. However, the in-depth analysis and evaluation of these ontologies with respect to quality-control metrics are missing. This work presents an overview of ontologies used in Materials Science and Engineering in order to guide the domain experts to decide which of the already available ontologies would be best suited for a particular purpose. We have identified 55 ontologies that are evaluated and assessed by adopting the well-known criteria of FAIR data principles [1]. Furthermore, statistical results about the reuse of these ontologies as well as essential ontology metrics are provided. Our analysis also provides information about whether the ontologies investigated are actively and openly maintained. With an overview of the identified ontologies, domain experts will be able to determine the intended ontologies and import and reuse terms from existing ontologies.</span></p>
<p><span>[1] Wilkinson, Mark D., et al. "The FAIR Guiding Principles for scientific data management and stewardship." </span><span>Scientific data</span><span> 3.1 (2016): 1-9.</span></p>
A Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals
Scholarly data is growing continuously containing information about the
articles from a plethora of venues including conferences, journals, etc. Many
initiatives have been taken to make scholarly data available as Knowledge
Graphs (KGs). These efforts to standardize these data and make them accessible
have also led to many challenges such as exploration of scholarly articles,
ambiguous authors, etc. This study more specifically targets the problem of
Author Name Disambiguation (AND) on Scholarly KGs and presents a novel
framework, Literally Author Name Disambiguation (LAND), which utilizes
Knowledge Graph Embeddings (KGEs) using multimodal literal information
generated from these KGs. This framework is based on three components: 1)
Multimodal KGEs, 2) A blocking procedure, and finally, 3) Hierarchical
Agglomerative Clustering. Extensive experiments have been conducted on two
newly created KGs: (i) KG containing information from Scientometrics Journal
from 1978 onwards (OC-782K), and (ii) a KG extracted from a well-known
benchmark for AND provided by AMiner (AMiner-534K). The results show that our
proposed architecture outperforms our baselines of 8-14% in terms of the F1
score and shows competitive performances on a challenging benchmark such as
AMiner. The code and the datasets are publicly available through Github:
https://github.com/sntcristian/and-kge and
Zenodo:https://doi.org/10.5281/zenodo.6309855 respectively
Towards a representation of temporal data in archival records: Use cases and requirements
Archival records are essential sources of information for historians and digital humanists to understand history. For modern information systems they are often analysed and integrated into Knowledge Graphs for better access, interoperability and re-use. However, due to restrictions of the representation of RDF predicates temporal data within archival records is a challenge to model. This position paper explains requirements for modeling temporal data in archival records based on running research projects in which archival records are analysed and integrated in Knowledge Graphs for research and exploration
Deep Learning meets Knowledge Graphs for Scholarly Data Classification
The amount of scientific literature continuously grows, which poses an increasing challenge for researchers to manage, find and explore research results. Therefore, the classification of scientific work is widely applied to enable the retrieval, support the search of suitable reviewers during the reviewing process, and in general to organize the existing literature according to a given schema. The automation of this classification process not only simplifies the submission process for authors, but also ensures the coherent assignment of classes. However, especially fine-grained classes and new research fields do not provide sufficient training data to automatize the process. Additionally, given the large number of not mutual exclusive classes, it is often difficult and computationally expensive to train models able to deal with multi-class multi-label settings. To overcome these issues, this work presents a preliminary Deep Learning framework as a solution for multi-label text classification for scholarly papers about Computer Science. The proposed model addresses the issue of insufficient data by utilizing the semantics of classes, which is explicitly provided by latent representations of class labels. This study uses Knowledge Graphs as a source of these required external class definitions by identifying corresponding entities in DBpedia to improve the overall classification
Linked Data Supported Content Analysis for Sociology
Philology and hermeneutics as the analysis and interpretation of natural language text in written historical sources are the predecessors of modern content analysis and date back already to antiquity. In empirical social sciences, especially in sociology, content analysis provides valuable insights to social structures and cultural norms of the present and past. With the ever growing amount of text on the web to analyze, also numerous computer-assisted text analysis techniques and tools were developed in sociological research. However, existing methods often go without sufficient standardization. As a consequence, sociological text analysis is lacking transparency, reproducibility and data re-usability.
The goal of this paper is to show, how Linked Data principles and Entity Linking techniques can be used to structure, publish and analyze natural language text for sociological research to tackle these shortcomings. This is achieved on the use case of constitutional text documents of the Netherlands from 1884 to 2016 which represent an important contribution to the European cultural heritage. Finally, the generated data is made available and re-usable as Linked Data not only for sociologists, but also for all other researchers in the digital humanities domain interested in the development of constitutions in the Netherlands
An assessment of deep learning models and word embeddings for toxicity detection within online textual comments
Today, increasing numbers of people are interacting online and a lot of textual comments are being produced due to the explosion of online communication. However, a paramount inconvenience within online environments is that comments that are shared within digital platforms can hide hazards, such as fake news, insults, harassment, and, more in general, comments that may hurt someone’s feelings. In this scenario, the detection of this kind of toxicity has an important role to moderate online communication. Deep learning technologies have recently delivered impressive performance within Natural Language Processing applications encompassing Sentiment Analysis and emotion detection across numerous datasets. Such models do not need any pre-defined hand-picked features, but they learn sophisticated features from the input datasets by themselves. In such a domain, word embeddings have been widely used as a way of representing words in Sentiment Analysis tasks, proving to be very effective. Therefore, in this paper, we investigated the use of deep learning and word embeddings to detect six different types of toxicity within online comments. In doing so, the most suitable deep learning layers and state-of-the-art word embeddings for identifying toxicity are evaluated. The results suggest that Long-Short Term Memory layers in combination with mimicked word embeddings are a good choice for this task
Recommended from our members
Contextual Language Models for Knowledge Graph Completion
Knowledge Graphs (KGs) have become the backbone of various
machine learning based applications over the past decade. However,
the KGs are often incomplete and inconsistent. Several representation
learning based approaches have been introduced to complete the missing
information in KGs. Besides, Neural Language Models (NLMs) have
gained huge momentum in NLP applications. However, exploiting the
contextual NLMs to tackle the Knowledge Graph Completion (KGC)
task is still an open research problem. In this paper, a GPT-2 based
KGC model is proposed and is evaluated on two benchmark datasets.
The initial results obtained from the _ne-tuning of the GPT-2 model
for triple classi_cation strengthens the importance of usage of NLMs for
KGC. Also, the impact of contextual language models for KGC has been
discussed
DDB-EDM to FaBiO: The case of the German Digital Library
Cultural heritage portals have the goal of providing users with seamless access to all their resources. This paper introduces initial e_orts for a user-oriented restructuring of the German Digital Library (DDB). At present, cultural heritage objects (CHOs) in the DDB are modeled using an extended version of the Europeana Data Model (DDBEDM), which negatively impacts usability and exploration. These challenges can be addressed by leveraging ontologies, and building a knowledge graph from the DDB's voluminous collection. Towards this goal, an alignment of bibliographic metadata from DDB-EDM to FRBR-Aligned Bibliographic Ontology (FaBiO) is presented
The Semantic Web. Latest Advances and New Domains - 13th International Conference, ESWC 2016
Challenges of applying knowledge graph and their embeddings to a real-world use-case
Different Knowledge Graph Embedding (KGE) models have been proposed so far which are trained on some specific KG completion tasks such as link prediction and evaluated on datasets which are mainly created for such purpose. Mostly, the embeddings learnt on link prediction tasks are not applied for downstream tasks in real-world use-cases such as data available in different companies/organizations. In this paper, the challenges with enriching a KG which is generated from a real-world relational database (RDB) about companies, with information from external sources such as Wikidata and learning representations for the KG are presented. Moreover, a comparative analysis is presented between the KGEs and various text embeddings on some downstream clustering tasks. The results of experiments indicate that in use-cases like the one used in this paper, where the KG is highly skewed, it is beneficial to use text embeddings or language models instead of KGEs
- …
