Proceedings of the International Conference on Dublin Core and Metadata Applications (DCMI)
Not a member yet
455 research outputs found
Sort by
Automatic Metadata Extraction From Museum Specimen Labels
This paper describes the information properties of museum
specimen labels and machine learning tools to automatically extract Darwin Core
(DwC) and other metadata from these labels processed through Optical Character
Recognition (OCR). The DwC is a metadata profile describing the core set of access
points for search and retrieval of natural history collections and observation
databases. Using the HERBIS Learning System (HLS) we extract 74 independent elements
from these labels. The automated text extraction tools are provided as a web service
so that users can reference digital images of specimens and receive back an extended
Darwin Core XML representation of the content of the label. This automated
extraction task is made more difficult by the high variability of museum label
formats, OCR errors and the open class nature of some elements. In this paper we
introduce our overall system architecture, and variability robust solutions
including, the application of Hidden Markov and Naïve Bayes machine learning models,
data cleaning, use of field element identifiers, and specialist learning models. The
techniques developed here could be adapted to any metadata extraction situation with
noisy text and weakly ordered elements
Exploring Evolutionary Biologists’ Use and Perceptions of Semantic Metadata for Data Curation
This poster will report on a study examining how evolutionary
biologists create and use personal metadata to organize their research data. Using
an ethnographic interview technique, participants are being interviewed about their
current and previous data organization styles and techniques. This information about
metadata and information organization can be used to inform new workflow and
organization models for knowledge organization and metadata creation practices in
developments for repositories, libraries, and cyberinfrastructures
Building a Terminology Network for Search: The KoMoHe Project
The paper reports about results on the GESIS-IZ project
“Competence Center Modeling and Treatment of Semantic Heterogeneity” (KoMoHe).
KoMoHe supervised a terminology mapping effort, in which ‘cross-concordances’
between major controlled vocabularies were organized, created and managed. In this
paper we describe the establishment and implementation of cross-concordances for
search in a digital library (DL)
Theme Creation for Digital Collections
This paper presents a solution to integrate multiple sources of
semantics for the purpose of metadata creation. A new framework is proposed to
define topics and themes with both manually and automatically generated terms. The
automatically generated terms include both terms from a semantic analysis of the
collections and terms from previous user’s queries. An interface is developed to
facilitate the creation of such topics and themes as well as the use of such topics
and themes to create metadata for digital resources. The framework and the interface
promote human-computer collaboration in metadata creation. Several principles
underlying such approach are also discussed
Relating Folksonomies with Dublin Core
Folksonomy is the result of describing the Web resources with
the use of tags created by the Web users. Although it has become a rich basis for
the description of resources, in general terms it is not beeing integrated in the
metadata. In order to be intelligible by machines and, therefore, used in the
Semantic Web context, they should be automatically allocated to specific metadata
elements. Thus, this paper presents part of a research carried out to continue the
project Kinds of Tags (KoT), which intends to identify elements of the metadata
originating from folksonomies and to propose an application profile for DC Social
Tagging. It will allow that the values reported by the tags may be conveniently
gathered by metadata interoperability protocols, such as the Open Archives
Initiative – Protocol for Metadata Harvesting (OAI-PMH). The results of the pilot
study that confirm some of the results of the KoT, show that a significant quantity
of tags could not be allocated to the Dublin Core Metadata Element Set (DCMES). New
properties, such as Action, Depth, Rate, and Utility are proposed. From the analysis
of every tag that is contained in the dataset, those potential new properties will
have to be validated by the DC Social Tagging Community
A Comparison of Social Tagging Designs and User Participation
Social tagging empowers users to categorize content in a
personally meaningful way while harnessing their potential to contribute to a
collaborative construction of knowledge. In addition, social tagging systems offer
innovative filtering mechanisms that facilitate resource discovery and browsing. As
a result, social tags support online communication, informal or intended learning as
well as the development of online communities. This poster will report on a mixed
methods study that examined how undergraduate students participate in social tagging
activities. Preliminary results of this study echo findings found in the growing
literature concerning social tagging from the fields of computer science and
information science
Encoding Application Profiles in a Computational Model of the Crosswalk
This paper describes the role of the Dublin Core Terms
application profile in the management of crosswalks involving MARC in OCLC’s
Crosswalk Web Service. This service, described in Godby, Smith and Childress (2008),
formalizes the notion of crosswalk (Getty, n.d.) by hiding technical details and
permitting the semantic equivalences to emerge as the centerpiece. As a result,
metadata experts, who are typically not programmers, can enter the translation logic
into a spreadsheet that can be automatically converted into executable code. The
Crosswalk Web Service supports many mappings involving standards for describing
bibliographic metadata, but the complex relationships among MARC and Dublin Core are
especially compelling because they would be far less elegantly managed without the
conceptual model of the application profile and the computational model of the
crosswalk. With its focus on elements that can be mixed, matched, added, and
redefined, the application profile (Heery and Patel, 2000) is a natural fit with the
translation model of the Crosswalk Web Service, which attempts to achieve
interoperability by mapping one pair of elements at a time. Users can test the
service with their own records by accessing the public demo on the OCLC
ResearchWorks page, or by invoking the Dublin Core export functions in OCLC’s
Connexion® Client
DCMF: DC & Microformats, a good marriage
This report introduces the Dublin Core Microformats (DCMF)
project, a new way to use the DC element set within X/HTML. The DC microformats
encode explicit semantic expressions in an X/HTML webpage, by using a specific list
of terms for values of the attributes “rev” and “rel” for "a" and "link" elements,
and “class” and “id” of other elements. Micrforomats can be easily processed by user
agents and software, enabling a high level of interoperability. These
characteristics are crucial for the growing number of social applications allowing
users to participate in the Web 2.0 environment as information creators and
consumers. This report reviews the origins of microformats; illustrates the coding
of DC microformats using the Dublin Core Metadata Gen tool, and a Firefox extension
for extraction and visualization; and discusses the benefits of creating Web
services utilizing DC microformats
Doing the LibraryThing in an Academic Library Catalog
What can an academic catalog look like in a Web 2.0
environment? This poster presents data analyzing the quality and quantity of the
metadata that a large academic library would expect to gain if utilizing a service
like that found on LibraryThing. It also looks at the advantaages and disadvantages
of controlled vocabularies and social tagging
Collection/Item Metadata Relationships
Contemporary retrieval systems, which search across
collections, usually ignore collection-level metadata. Alternative approaches,
exploiting collection-level information, will require an understanding of the
various kinds of relationships that can obtain between collection-level and
item-level metadata. This paper outlines the problem and describes a project that is
developing a logic-based framework for classifying collection/item metadata
relationships. This framework will support (i) metadata specification developers
defining metadata elements, (ii) metadata creators describing objects, and (iii)
system designers implementing systems that take advantage of collection-level
metadata. We present two simple examples of collection/item metadata relationship
categories, attribute/value-propagation and value-propagation, and show that even in
these simple cases a precise formulation requires modal notions in addition to
first-order logic. These formulations are related to recent work in information
retrieval and ontology evaluation