International Journal of Digital Curation
Not a member yet
    605 research outputs found

    A Short Story about XML Schemas, Digital Preservation and Format Libraries

    Get PDF
    One morning we came in to work to find that one of our servers had made 1.5 million attempts to contact an external server in the preceding hour. It turned out that the calls were being generated by the Library’s digital preservation system (Rosetta) while attempting to validate XML Schema Definition (XSD) declarations included in the XML files of the Library’s online newspaper application Papers Past, which we were in the process of loading into Rosetta. This paper describes our response to this situation and outlines some of the issues that needed to be canvassed before we were able to arrive at a suitable solution, including the digital preservation status of these XSDs; their impact on validation tools, such as JHOVE; and where these objects should reside if they are considered material to the digital preservation process

    Grammar-Based Specification and Parsing of Binary File Formats

    Get PDF
    The capability to validate and view or play binary file formats, as well as to convert binary file formats to standard or current file formats, is critically important to the preservation of digital data and records. This paper describes the extension of context-free grammars from strings to binary files. Binary files are arrays of data types, such as long and short integers, floating-point numbers and pointers, as well as characters. The concept of an attribute grammar is extended to these context-free array grammars. This attribute grammar has been used to define a number of chunk-based and directory-based binary file formats. A parser generator has been used with some of these grammars to generate syntax checkers (recognizers) for validating binary file formats. Among the potential benefits of an attribute grammar-based approach to specification and parsing of binary file formats is that attribute grammars not only support format validation, but support generation of error messages during validation of format, validation of semantic constraints, attribute value extraction (characterization), generation of viewers or players for file formats, and conversion to current or standard file formats. The significance of these results is that with these extensions to core computer science concepts, traditional parser/compiler technologies can potentially be used as a part of a general, cost effective curation strategy for binary file formats

    The Data Management Skills Support Initiative: Synthesising Postgraduate Training in Research Data Management

    Get PDF
    This paper will describe the efforts and findings of the JISC Data Management Skills Support Initiative (‘DaMSSI’). DaMSSI was co-funded by the JISC Managing Research Data programme and the Research Information Network (RIN), in partnership with the Digital Curation Centre, to review, synthesise and augment the training offerings of the JISC Research Data Management Training Materials (‘RDMTrain’) projects.DaMSSI tested the effectiveness of the Society of College, National and University Libraries’ Seven Pillars of Information Literacy model (SCONUL, 2011), and Vitae’s Researcher Development Framework (‘Vitae RDF’) for consistently describing research data management (‘RDM’) skills and skills development paths in UK HEI postgraduate courses.With the collaboration of the RDMTrain projects, we mapped individual course modules to these two models and identified basic generic data management skills alongside discipline-specific requirements. A synthesis of the training outputs of the projects was then carried out, which further investigated the generic versus discipline-specific considerations and other successful approaches to training that had been identified as a result of the projects’ work. In addition we produced a series of career profiles to help illustrate the fact that data management is an essential component – in obvious and not-so-obvious ways – of a wide range of professions.We found that both models had potential for consistently and coherently describing data management skills training and embedding this within broader institutional postgraduate curricula. However, we feel that additional discipline-specific references to data management skills could also be beneficial for effective use of these models. Our synthesis work identified that the majority of core skills were generic across disciplines at the postgraduate level, with the discipline-specific approach showing its value in engaging the audience and providing context for the generic principles.Findings were fed back to SCONUL and Vitae to help in the refinement of their respective models, and we are working with a number of other projects, such as the DCC and the EC-funded Digital Curator Vocational Education Europe (DigCurV2) initiative, to investigate ways to take forward the training profiling work we have begun

    Assessing Migration Risk for Scientific Data Formats

    Get PDF
    The majority of information about science, culture, society, economy and the environment is born digital, yet the underlying technology is subject to rapid obsolescence. One solution to this obsolescence, format migration, is widely practiced and supported by many software packages, yet migration has well known risks. For example, newer formats – even where similar in function – do not generally support all of the features of their predecessors, and, where similar features exist, there may be significant differences of interpretation.There appears to be a conflict between the wide use of migration and its known risks. In this paper we explore a simple hypothesis – that, where migration paths exist, the majority of data files can be safely migrated leaving only a few that must be handled more carefully – in the context of several scientific data formats that are or were widely used. Our approach is to gather information about potential migration mismatches and, using custom tools, evaluate a large collection of data files for the incidence of these risks. Our results support our initial hypothesis, though with some caveats. Further, we found that writing a tool to identify “risky” format features is considerably easier than writing a migration tool

    Towards the Development of a Test Corpus of Digital Objects for the Evaluation of File Format Identification Tools and Signatures

    Get PDF
    The digital preservation community currently utilises a number of tools and automated processes to identify and validate digital objects. The identification of digital objects is a vital first step in their long-term preservation, but the results returned by tools used for this purpose are lacking in transparency, and are not easily tested or verified. This paper suggests that a test corpus of digital objects is one way of providing this verification and validation, ultimately improving trust in the tools, and providing further stimulus to their development. Issues to be considered are outlined, and attention is drawn to particular examples of existing digital corpora which could conceivably provide a useable framework or starting point for our own communities needs. This paper does not seek to answer all questions in this area, but merely attempts to set out areas for consideration in any next step that is taken

    Peer-Reviewed Open Research Data: Results of a Pilot

    Get PDF
    Peer review of publications is at the core of science and primarily seen as instrument for ensuring research quality. However, it is less common to independently value the quality of the underlying data as well. In the light of the ‘data deluge’ it makes sense to extend peer review to the data itself and this way evaluate the degree to which the data are fit for re-use. This paper describes a pilot study at EASY - the electronic archive for (open) research data at our institution. In EASY, researchers can archive their data and add metadata themselves. Devoted to open access and data sharing, at the archive we are interested in further enriching these metadata with peer reviews.As a pilot, we established a workflow where researchers who have downloaded data sets from the archive were asked to review the downloaded data set. This paper describes the details of the pilot including the findings, both quantitative and qualitative. Finally, we discuss issues that need to be solved when such a pilot is turned into a structural peer review functionality for the archiving system

    Development of a Pilot Data Management Infrastructure for Biomedical Researchers at University of Manchester – Approach, Findings, Challenges and Outlook of the MaDAM Project

    Get PDF
    Management and curation of digital data has been becoming ever more important in a higher education and research environment characterised by large and complex data, demand for more interdisciplinary and collaborative work, extended funder requirements and use of e-infrastructures to facilitate new research methods and paradigms. This paper presents the approach, technical infrastructure, findings, challenges and outlook (including future development within the successor project, MiSS) of the ‘MaDAM: Pilot data management infrastructure for biomedical researchers at University of Manchester’ project funded under the infrastructure strand of the JISC Managing Research Data (JISCMRD) programme. MaDAM developed a pilot research data management solution at the University of Manchester based on biomedical researchers’ requirements, which includes technical and governance components with the flexibility to meet future needs across multiple research groups and disciplines

    Scientific Research Data Management for Soil-Vegetation-Atmosphere Data – The TR32DB

    Get PDF
    The implementation of a scientific research data management system is an important task within long-term, interdisciplinary research projects. Besides sustainable storage of data, including accurate descriptions with metadata, easy and secure exchange and provision of data is necessary, as well as backup and visualisation. The design of such a system poses challenges and problems that need to be solved.This paper describes the practical experiences gained by the implementation of a scientific research data management system, established in a large, interdisciplinary research project with focus on Soil-Vegetation-Atmosphere Data

    Making Data a First Class Scientific Output: Data Citation and Publication by NERC’s Environmental Data Centres

    Get PDF
    The NERC Science Information Strategy Data Citation and Publication project aims to develop and formalise a method for formally citing and publishing the datasets stored in its environmental data centres. It is believed that this will act as an incentive for scientists, who often invest a great deal of effort in creating datasets, to submit their data to a suitable data repository where it can properly be archived and curated. Data citation and publication will also provide a mechanism for data producers to receive credit for their work, thereby encouraging them to share their data more freely

    Developing Virtual CD-ROM Collections: The Voyager Company Publications

    Get PDF
    Over the past 20 years, many thousands of CD-ROM titles were published; many of these have lasting cultural significance, yet present a difficult challenge for libraries due to obsolescence of the supporting software and hardware, and the consequent decline in the technical knowledge required to support them. The current trend appears to be one of abandonment – for example, the Indiana University Libraries no longer maintain machines capable of accessing early CD-ROM titles.In previous work, we proposed an access model based upon networked ‘virtual collections’ of CD-ROMs which can enable consortia of libraries to pool the technical expertise necessary to provide continued access to such materials for a geographically sparse base of patrons, who may have limited technical knowledge.In this paper, we extend this idea to CD-ROMs designed to operate on ‘classic’ Macintosh systems with an extensive case study – the catalog of the Voyager Company publications, which was the first major innovator in interactive CD-ROMs. The work described includes emulator extensions to support obsolete CD formats and to enable networked access to the virtual collection

    522

    full texts

    605

    metadata records
    Updated in last 30 days.
    International Journal of Digital Curation
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇