International Journal of Digital Curation
Not a member yet
    605 research outputs found

    Towards Smart Storage for Repository Preservation Services

    Get PDF
    The move to digital is being accompanied by a huge rise in volumes of (born-digital) content and data. As a result the curation lifecycle has to be redrawn. Processes such as selection and evaluation for preservation have to be driven by automation. Manual processes will not scale, and the traditional signifiers and selection criteria in older formats, such as print publication, are changing. The paper will examine at a conceptual and practical level how preservation intelligence can be built into software-based digital preservation tools and services on the Web and across the network ‘cloud’ to create ‘smart’ storage for long-term, continuous data monitoring and management. Some early examples will be presented, focussing on storage management and format risk assessment

    Editorial

    Get PDF
    Chris Rusbridge and Kevin Ashley introduce Issue 1, Volume

    A Framework for Distributed Preservation Workflows

    Get PDF
    The Planets Project is developing a service-oriented environment for the definition and evaluation of preservation strategies for human-centric data. It focuses on the question of logically preserving digital materials, as opposed to the physical preservation of content bit-streams. This includes the development of preservation tools for the automated characterisation, migration, and comparison of different types of Digital Objects as well as the emulation of their original runtime environment in order to ensure long-time access and interpretability. The Planets integrated environment provides a number of end-user applications that allow data curators to execute and scientifically evaluate preservation experiments based on composable preservation services. In this paper, we focus on the middleware and programming model and show how it can be utilised in order to create complex preservation workflows

    Relay-supporting Archives: Requirements and Progress

    Get PDF
    We characterize long-term preservation of digital content as an extended relay in time, in which repeated handoffs of information occur independently at every architectural layer: at the physical layer, where bits are handed off between storage systems; at the logical layer, where digital objects are handed off between repository systems; and at the administrative layer, where collections of objects and relationships are handed off between archives, curators, and institutions.  We examine the support of current preservation technologies for these handoffs, note shortcomings, and argue that some modest improvements would result in a "relay-supporting" preservation infrastructure, one that provides a baseline level of preservation by mitigating the risk of fundamental information loss.  Finally, we propose a series of tests to validate a relay-supporting infrastructure, including a second Archive Ingest and Handling Test (AIHT)

    Embedding Metadata and Other Semantics in Word Processing Documents

    Get PDF
    This paper describes a technique for embedding document metadata, and potentially other semantic references inline in word processing documents, which the authors have implemented with the help of a software development team. Several assumptions underly the approach; It must be available across computing platforms and work with both Microsoft Word (because of its user base) and OpenOffice.org (because of its free availability). Further the application needs to be acceptable to and usable by users, so the initial implementation covers only small number of features, which will only be extended after user-testing. Within these constraints the system provides a mechanism for encoding not only simple metadata, but for inferring hierarchical relationships between metadata elements from a ‘flat’ word processing file.The paper includes links to open source code implementing the techniques as part of a broader suite of tools for academic writing. This addresses tools and software, semantic web and data curation, integrating curation into research workflows and will provide a platform for integrating work on ontologies, vocabularies and folksonomies into word processing tools

    Modelling Organizational Preservation Goals to Guide Digital Preservation

    Get PDF
    This paper is an extended and updated version of the work reported at iPres 2008. Digital preservation activities can only succeed if they go beyond the technical properties of digital objects. They must consider the strategy, policy, goals, and constraints of the institution that undertakes them and take into account the cultural and institutional framework in which data, documents and records are preserved. Furthermore, because organizations differ in many ways, a one-size-fits-all approach cannot be appropriate. Fortunately, organizations involved in digital preservation have created documents describing their policies, strategies, work-flows, plans, and goals to provide guidance. They also have skilled staff who are aware of sometimes unwritten considerations. Within Planets (Farquhar & Hockx-Yu, 2007), a four-year project co-funded by the European Union to address core digital preservation challenges, we have analyzed preservation guiding documents and interviewed staff from libraries, archives, and data centres that are actively engaged in digital preservation. This paper introduces a conceptual model for expressing the core concepts and requirements that appear in preservation guiding documents. It defines a specific vocabulary that institutions can reuse for expressing their own policies and strategies. In addition to providing a conceptual framework, the model and vocabulary support automated preservation planning tools through an XML representation

    Data Curation Program Development in U.S. Universities: The Georgia Institute of Technology Example

    Get PDF
    The curation of scientific research data at U.S. universities is a story of enterprising individuals and of incremental progress. A small number of libraries and data centers who see the possibilities of becoming “digital information management centers” are taking entrepreneurial steps to extend beyond their traditional information assets and include managing scientific and scholarly research data. The Georgia Institute of Technology (GT) has had a similar development path toward a data curation program based in its library. This paper will articulate GT’s program development, which the author offers as an experience common in U.S. universities. The main characteristic is a program devoid of top-level mandates and incentives, but rich with independent, “bottom-up” action. The paper will address program antecedents and context, inter-institutional partnerships that advance the library’s curation program, library organizational developments, partnerships with campus research communities, and a proposed model for curation program development. It concludes that despite the clear need for data curation put forth by researchers such as the groups of neuroscientists and bioscientists referenced in this paper, the university experience examined suggests that gathering resources for developing data curation programs at the institutional level is proving to be a quite onerous. However, and in spite of the challenges, some U.S. research universities are beginning to establish perceptible data curation programs

    Using the DCC Lifecycle Model to Curate a Gene Expression Database: A Case Study

    Get PDF
    Developmental Gene Expression Map (DGEMap) is an EU-funded Design Study, which will accelerate an integrated European approach to gene expression in early human development. As part of this design study, we have had to address the challenges and issues raised by the long-term curation of such a resource. As this project is primarily one of data creators, learning about curation, we have been looking at some of the models and tools that are already available in the digital curation field in order to inform our thinking on how we should proceed with curating DGEMap. This has led us to uncover a wide range of resources for data creators and curators alike. Here we will discuss the future curation of DGEMap as a case study. We believe our experience could be instructive to other projects looking to improve the curation and management of their data

    Curating the CIA World Factbook

    Get PDF
    The CIA World Factbook is a prime example of a curated database – a database that is constructed and maintained with a great deal of human effort in collecting, verifying, and annotating data. Preservation of old versions of the Factbook is important for verification of citations; it is also essential for anyone interested in the history of the data such as demographic change. Although the Factbook has been published, both physically and electronically, only for the past 30 years, we appear in danger of losing this history. This paper investigates the issues involved in capturing the history of an evolving database and its application to the CIA World Factbook. In particular it shows that there is substantial added value to be gained by preserving databases in such a way that questions about the change in data, (longitudinal queries) can be readily answered. Within this paper, we describe techniques for recording change in a curated database and we describe novel techniques for querying the change. Using the example of this archived curated database, we discuss the extent to which the accepted practices and terminology of archiving, curation and digital preservation apply to this important class of digital artefacts

    Long-term Preservation of Earth Observation Data and Knowledge in ESA through CASPAR

    Get PDF
    ESA-ESRIN, the European Space Agency Centre for Earth Observation (EO), is the largest European EO data provider and operates as the reference European centre for EO payload data exploitation. EO Space Missions provide global coverage of the Earth across both space and time generating on a routine continuous basis huge amounts of data (from a variety of sensors) that need to be acquired, processed, elaborated, appraised and archived by dedicated systems. Long-term Preservation of these data and of the ability to discover, access and process them is a fundamental issue and a major challenge at programmatic, technological and operational levels.Moreover these data are essential for scientists needing broad series of data covering long time periods and from many sources. They are used for many types of investigations including ones of international importance such as the study of the Global Change and the Global Monitoring for Environment and Security (GMES) Program. Therefore it is of primary importance not only to guarantee easy accessibility of historical data but also to ensure users are able to understand and use them; in fact data interpretation can be even more complicated given the fact that scientists may not have (or may not have access to) the right knowledge to interpret these data correctly.To satisfy these requirements, the European Space Agency (ESA), in addition to other internal initiatives, is participating in several EU-funded projects such as CASPAR (Cultural, Artistic, and Scientific knowledge for Preservation, Access and Retrieval), which is building a framework to support the end-to-end preservation lifecycle for digital information, based on the OAIS reference model, with a strong focus on the preservation of the knowledge associated with data.In the CASPAR Project ESA plays the role of both user and infrastructure provider for one of the scientific testbeds, putting into effect dedicated scenarios with the aim of validating the CASPAR solutions in the Earth Science domain. The other testbeds are in the domains of Cultural Heritage and of Contemporary Performing Arts; together they provide a severe test of preservation tools and techniques.In the context of the current ESA overall strategies carried out in collaboration with European EO data owners/providers, entities and institutions which have the objective of guaranteeing long-term preservation of EO data and knowledge, this paper will focus on the ESA participation and contribution to the CASPAR Project, describing in detail the implementation of the ESA scientific testbed

    522

    full texts

    605

    metadata records
    Updated in last 30 days.
    International Journal of Digital Curation
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇