International Journal of Digital Curation
Not a member yet
605 research outputs found
Sort by
Migrating Home Computer Audio Waveforms to Digital Objects: A Case Study on Digital Archaeology
Rescuing data from inaccessible or damaged storage media for the purpose of preserving the digital data for the long term is one of the dimensions of digital archaeology. With the current pace of technological development, any system can become obsolete in a matter of years and hence the data stored in a specific storage media might not be accessible anymore due to the unavailability of the system to access the media. In order to preserve digital records residing in such storage media, it is necessary to extract the data stored in those media by some means.One early storage medium for home computers in the 1980s was audio tape. The first home computer systems allowed the use of standard cassette players to record and replay data. Audio cassettes are more durable than old home computers when properly stored. Devices playing this medium (i.e. tape recorders) can be found in working condition or can be repaired, as they are usually made out of standard components. By re-engineering the format of the waveform and the file formats, the data on such media can then be extracted from a digitised audio stream and migrated to a non-obsolete format.In this paper we present a case study on extracting the data stored on an audio tape by an early home computer system, namely the Philips Videopac+ G7400. The original data formats were re-engineered and an application was written to support the migration of the data stored on tapes without using the original system. This eliminates the necessity of keeping an obsolete system alive for enabling access to the data on the storage media meant for this system. Two different methods to interpret the data and eliminate possible errors in the tape were implemented and evaluated on original tapes, which were recorded 20 years ago. Results show that with some error correction methods, parts of the tapes are still readable even without the original system. It also implies that it is easier to build solutions while original systems are still available in a working condition
Dependency Analysis of Legacy Digital Materials to Support Emulation Based Preservation
Emulation has been widely discussed as a preservation strategy for digital documents that depend upon proprietary executables, as well as for legacy programs. The fundamental assumption of this strategy is that an artifact (document or program) will be bundled with any required contemporaneous software in an archival information package (AIP) which can be loaded and executed in an emulation environment by patrons wishing to access the preserved artifact, yet little has been written about how to identify the required components for such an AIP. Even where a digital document was distributed with a binary viewer, there may be dependencies on other software libraries. In this paper we discuss a pilot study that performed dependency analysis for digital materials originally distributed on CD-ROM. In particular, we show how to utilize a small number of existing off-the-shelf libraries to build a tool that can analyze executables within ISO (CD-ROM) images, and then examine the results of applying this tool to a body of archived images
What Constitutes Successful Format Conversion? Towards a Formalization of \u27Intellectual Content\u27
Recent work in the semantics of markup languages may offer a way to achieve more reliable results for format conversion, or at least a way to state the goal more explicitly. In the work discussed, the meaning of markup in a document is taken as the set of things accepted as true because of the markup\u27s presence, or equivalently, as the set of inferences licensed by the markup in the document. It is possible, in principle, to apply a general semantic description of a markup vocabulary to documents encoded using that vocabulary and to generate a set of inferences (typically rather large, but finite) as a result. An ideal format conversion translating a digital object from one vocabulary to another, then, can be characterized as one which neither adds nor drops any licensed inferences; it is possible to check this equivalence explicitly for a given conversion of a digital object, and possible in principle (although probably beyond current capabilities in practice) to prove that a given transformation will, if given valid and semantically correct input, always produce output that is semantically equivalent to its input. This approach is directly applicable to the XML formats frequently used for scientific and other data, but it is also easily generalized from SGML/XML-based markup languages to digital formats in general; at a high level, it is equally applicable to document markup, to database exchanges, and to ad hoc formats for high-volume scientific data.Some obvious complications and technical difficulties arising from this approach are discussed, as are some important implications. In most real-world format conversions, the source and target formats differ at least somewhat in their ontology, either in the level of detail they cover or in the way they carve reality into classes; it is thus desirable not only to define what a perfect format conversion looks like, but to quantify the loss or distortion of information resulting from the conversion
Digital Curation: The Emergence of a New Discipline
In the mid 1990s UK digital preservation activity concentrated on ensuring the survival of digital material – spurred on by the US report Preserving Digital Information (The Task Force on Archiving of Digital Information, 1996) and developed through JISC-funded activities. Technical developments and a maturing understanding of organisational activity and workflow saw the emphasis move to ensuring the access, use and reuse of digital materials throughout their lifecycle. Digital Curation emerged as a new discipline supported through the activities of the UK’s Digital Curation Centre and a number of EU 6th Framework Projects. Digital Curation is now embedded in both practice and research; with the development of tools, and the foundation of a number of support units and academic educators offering training and furthering research
DataStaR: Using the Semantic Web approach for Data Curation
In disciplines as varied as medicine, social sciences, and economics, data and their analyses are essential parts of researchers’ contributions to their respective fields. While sharing research data for review and analysis presents new opportunities for furthering research, capturing these data in digital forms and providing the digital infrastructure for sharing data and metadata pose several challenges. This paper reviews the motivations behind and design of the Data Staging Repository (DataStaR) platform that targets specific portions of the research data curation lifecycle: data and metadata capture and sharing prior to publication, and publication to permanent archival repositories. The goal of DataStaR is to support both the sharing and publishing of data while at the same time enabling metadata creation without imposing additional overheads for researchers and librarians. Furthermore, DataStaR is intended to provide cross-disciplinary support by being able to integrate different domain-specific metadata schemas according to researchers’ needs. DataStaR’s strategy of a usable interface coupled with metadata flexibility allows for a more scaleable solution for data sharing, publication, and metadata reuse
Curating Complex, Dynamic and Distributed Data: Telehealth as a Laboratory for Strategy
Telehealth monitoring data is now being collected across large populations of patients with chronic diseases such as stroke, hypertension, COPD and dementia. These large, complex and heterogeneous datasets, including distributed sensor and mobile datasets, present real opportunities for knowledge discovery and re-use, however they also generate new challenges for curation. This paper uses qualitative research with stakeholders in two nationally-funded telehealth projects to outline the perceptions, practices and preferences of different stakeholders with regard to data curation. Telehealth provides a living laboratory for the very different challenges implicit in designing and managing data infrastructure for embedded and ubiquitous computing. Here, technical and human agents are distributed, and interaction and state change is a central component of design, rather than an inconvenient challenge to it. The authors argue that there are lessons to be learned from other domains where data infrastructure has been radically rethought to address these challenges
Born Broken: Fonts and Information Loss in Legacy Digital Documents
For millions of legacy documents, correct rendering depends upon resources such as fonts that are not generally embedded within the document structure. Yet there is a significant risk of information loss due to missing or incorrectly substituted fonts. Large document collections depend on thousands of unique fonts not available on a common desktop workstation, which typically has between 100 and 200 fonts. Silent substitution of fonts, performed by applications such as Microsoft Office, can yield poorly rendered documents. In this paper we use a collection of 230,000 Word documents to assess the difficulty of matching font requirements with a database of fonts. We describe the identifying information contained in common font formats, font requirements stored in Word documents, the API provided by Windows to support font requests by applications, the documented substitution algorithms used by Windows when requested fonts are not available, and the ways in which support software might be used to control font substitution in a preservation environment
Towards Support for Long-Term Digital Preservation in Product Life Cycle Management
Important legal and economic motivations exist for the design and engineering industry to address and integrate digital long-term preservation into product life cycle management (PLM). Investigations revealed that it is not sufficient to archive only the product design data which is created in early PLM phases, but preservation is needed for data that is produced during the entire product lifecycle including early and late phases. Data that is relevant for preservation consists of requirements analysis documents, design rationale, data that reflects experiences during product operation and also metadata like social collaboration context. In addition, also the engineering environment itself that contains specific versions of all tools and services is a candidate for preservation. This paper takes a closer look at engineering preservation use case scenarios as well as PLM characteristics and workflows that are relevant for long-term preservation. Resulting requirements for a long-term preservation system lead to an OAIS (Open Archival Information System) based system architecture and a proposed preservation service interface that respects the needs of the engineering industry
Educating Digital Curators: Challenges and Opportunities
This paper describes a number of critical challenges faced by digital curation educators and suggests how the choices we make in building educational programs may impact the development of curation as a professional discipline. We focus on curriculum and program building as key steps in defining the educational needs of curators, and we argue for greater collaboration among educators, researchers and practitioners in the field, as a way to speed the emergence of curation as a discipline and to foster the integration of curation programs within libraries and archives
Research Data Management Initiatives at University of Edinburgh
During the last decade, national and international attention has been increasingly focused on issues of research data management and access to publicly funded research data. The pressure brought to bear on researchers to improve their data management and data sharing practice has come from research funders seeking to add value to expensive research and solve cross-disciplinary grand challenges; publishers seeking to be responsive to calls for transparency and reproducibility of the scientific record; and the public seeking to gain and re-use knowledge for their own purposes using new online tools. Meanwhile higher education institutions have been rather reluctant to assert their role in either incentivising or supporting their academic staff in meeting these more demanding requirements for research practice, partly due to lack of knowledge as to how to provide suitable assistance or facilities for data storage and curation/preservation. This paper discusses the activities and drivers behind one institution’s recent attempts to address this gap, with reflection on lessons learned and future direction