International Journal of Digital Curation
Not a member yet
605 research outputs found
Sort by
(Not) Playing Games: Player-Produced Walkthroughs as Archival Documents of Digital Gameplay
The subject of digital game preservation is one that has moved up the research agenda in recent years with a number of international projects, such as KEEP and Preserving Virtual Worlds, highlighting and seeking to address the impact of media decay, hardware and software obsolescence through different strategies including code emulation, for instance. Similarly, and reflecting a popular interest in the histories of digital games, exhibitions such as Game On (Barbican, UK) and GameCity (Nottingham, UK) experiment with ways of presenting games to a general audience. This article focuses on the UK’s National Videogame Archive (NVA) which, since its foundation in 2008, has developed approaches that both dovetail with and critique existing strategies to game preservation, exhibition and display.The article begins by noting the NVA’s interest in preserving not only the code or text of the game, but also the experience of using it – that is, the preservation of gameplay as well as games. This approach is born of a conceptualisation of digital games as what Moulthrop (2004) has called “configurative performances” that are made through the interaction of code, systems, rules and, essentially, the actions of players at play. The analysis develops by problematising technical solutions to game preservation by exploring the way seemingly minute differences in code execution greatly impact on this user experience.Given these issues, the article demonstrates how the NVA returns to first principles and questions the taken-for-granted assumption that the playable game is the most effective tool for interpretation. It also encourages a consideration of the uses of non-interactive audiovisual and (para)textual materials in game preservation activity. In particular, the focus falls upon player-produced walkthrough texts, which are presented as archetypical archival documents of gameplay. The article concludes by provocatively positing that these non-playable, non-interactive texts might be more useful to future game scholars than the playable game itself
Use of Ontologies for Data Integration and Curation
Data curation includes the goal of facilitating the re-use and combination of datasets, which is often impeded by incompatible data schema. Can we use ontologies to help with data integration? We suggest a semi-automatic process that involves the use of automatic text searching to help identify overlaps in metadata that accompany data schemas, plus human validation of suggested data matches.Problems include different text used to describe the same concept, different forms of data recording and different organizations of data. Ontologies can help by focussing attention on important words, providing synonyms to assist matching, and indicating in what context words are used. Beyond ontologies, data on the statistical behavior of data can be used to decide which data elements appear to be compatible with which other data elements. When curating data which may have hundreds or even thousands of data labels, semi-automatic assistance with data fusion should be of great help
Automation of Flexible Migration Workflows
Many digital preservation scenarios are based on the migration strategy, which itself is heavily tool-dependent. For popular, well-defined and often open file formats – e.g., digital images, such as PNG, GIF, JPEG – a wide range of tools exist. Migration workflows become more difficult with proprietary formats, as used by the several text processing applications becoming available in the last two decades. If a certain file format can not be rendered with actual software, emulation of the original environment remains a valid option. For instance, with the original Lotus AmiPro or Word Perfect, it is not a problem to save an object of this type in ASCII text or Rich Text Format. In specific environments, it is even possible to send the file to a virtual printer, thereby producing a PDF as a migration output. Such manual migration tasks typically involve human interaction, which may be feasible for a small number of objects, but not for larger batches of files.We propose a novel approach using a software-operated VNC abstraction layer in order to replace humans with machine interaction. Emulators or virtualization tools equipped with a VNC interface are very well suited for this approach. But screen, keyboard and mouse interaction is just part of the setup. Furthermore, digital objects need to be transferred into the original environment in order to be extracted after processing. Nevertheless, the complexity of the new generation of migration services is quickly rising; a preservation workflow is now comprised not only of the migration tool itself, but of a complete software and virtual hardware stack with recorded workflows linked to every supported migration scenario. Thus the requirements of OAIS management must include proper software archiving, emulator selection, system image and recording handling. The concept of view-paths could help either to automatically determine the proper pre-configured virtual environment or to set up system images for certain migration workflows. View-paths may rise in demand, as the generation of PDF output files from Word Perfect input could be cached as pre-fabricated emulator system images. The current groundwork provides several possible optimizations, such as using the automation features of the original environments
Linking to Scientific Data: Identity Problems of Unruly and Poorly Bounded Digital Objects
Within information systems, a significant aspect of search and retrieval across information objects, such as datasets, journal articles, or images, relies on the identity construction of the objects. This paper uses identity to refer to the qualities or characteristics of an information object that make it definable and recognizable, and can be used to distinguish it from other objects. Identity, in this context, can be seen as the foundation from which citations, metadata and identifiers are constructed.In recent years the idea of including datasets within the scientific record has been gaining significant momentum, with publishers, granting agencies and libraries engaging with the challenge. However, the task has been fraught with questions of best practice for establishing this infrastructure, especially in regards to how citations, metadata and identifiers should be constructed. These questions suggests a problem with how dataset identities are formed, such that an engagement with the definition of datasets as conceptual objects is warranted.This paper explores some of the ways in which scientific data is an unruly and poorly bounded object, and goes on to propose that in order for datasets to fulfill the roles expected for them, the following identity functions are essential for scholarly publications: (i) the dataset is constructed as a semantically and logically concrete object, (ii) the identity of the dataset is embedded, inherent and/or inseparable, (iii) the identity embodies a framework of authorship, rights and limitations, and (iv) the identity translates into an actionable mechanism for retrieval or reference
Cost Model for Digital Preservation: Cost of Digital Migration
The Danish Ministry of Culture has funded a project to set up a model for costing preservation of digital materials held by national cultural heritage institutions. The overall objective of the project was to increase cost effectiveness of digital preservation activities and to provide a basis for comparing and estimating future cost requirements for digital preservation. In this study we describe an activity-based costing methodology for digital preservation based on the Open Archice Information System (OAIS) Reference Model. Within this framework, which we denote the Cost Model for Digital Preservation (CMDP), the focus is on costing the functional entity Preservation Planning from the OAIS and digital migration activities. In order to estimate these costs we have identified cost-critical activities by analysing the functions in the OAIS model and the flows between them. The analysis has been supplemented with findings from the literature, and our own knowledge and experience. The identified cost-critical activities have subsequently been deconstructed into measurable components, cost dependencies have been examined, and the resulting equations expressed in a spreadsheet. Currently the model can calculate the cost of different migration scenarios for a series of preservation formats for text, images, sound, video, geodata, and spreadsheets. In order to verify the model it has been tested on cost data from two different migration projects at the Danish National Archives (DNA). The study found that the OAIS model provides a sound overall framework for the cost breakdown, but that some functions need additional detailing in order to cost activities accurately. Running the two sets of empirical data showed among other things that the model underestimates the cost of manpower-intensive migration projects, while it reinstates an often underestimated cost, which is the cost of developing migration software. The model has proven useful for estimating the costs of preservation planning and digital migrations. However, more work is needed to refine the existing equations and include the other functional entities of the OAIS model. Also the user-friendliness of the spreadsheet tool must be improved in future versions of the model. The CMDP is presently closing its second phase, where it has been extended to include the OAIS Functional Entity Ingest. This has also enabled us to adjust the theoretical model further, especially regarding the accuracy and precision of the model and in relation to the underlying parameters used in the equations, such as migration frequency and format complexity. Understanding the nature of digital preservation cost is prerequisite for increasing the overall efficiency, and achieving first quality for preservation of cultural heritage materials
Automating the Extraction of Metadata from Archaeological Data Using iRods Rules
The Texas Advanced Computing Center and the Institute for Classical Archaeology at the University of Texas at Austin developed a method that uses iRods rules and a Jython script to automate the extraction of metadata from digital archaeological data. The first step was to create a record-keeping system to classify the data. The record-keeping system employs file and directory hierarchy naming conventions designed specifically to maintain the relationship between the data objects and map the archaeological documentation process. The metadata implicit in the record-keeping system is automatically extracted upon ingest, combined with additional sources of metadata, and stored alongside the data in the iRods preservation environment. This method enables a more organized workflow for the researchers, helps them archive their data close to the moment of data creation, and avoids error prone manual metadata input. We describe the types of metadata extracted and provide technical details of the extraction process and storage of the data and metadata
Education for eScience Professionals: Integrating Data Curation and Cyberinfrastructure
Large, collaboratively managed datasets have become essential to many scientific and engineering endeavors, and their management has increased the need for "eScience professionals" who solve large scale information management problems for researchers and engineers. This paper considers the dimensions of work, worker, and workplace, including the knowledge, skills, and abilities needed for eScience professionals. We used focus groups and interviews to explore the needs of scientific researchers and how these needs may translate into curricular and program development choices. A cohort of five masters students also worked in targeted internship settings and completed internship logs. We organized this evidence into a job analysis that can be used for curriculum and program development at schools of library and information science
Are you Ready? Assessing Whether Organisations are Prepared for Digital Preservation
In early 2009 the Planets project undertook a survey of national libraries, archives, and other content-holding organisations in Europe to better understand the organisations\u27 digital preservation activities and needs, and to ensure that Planets\u27 technology and services are designed to meet them. Over 200 responses were received including a cross-section of major libraries and archives especially in Europe. The results provide a snapshot of organisations\u27 readiness to preserve digital collections for the future. The survey revealed a high level of awareness of the challenges of digital preservation within organisations. Findings indicated that approximately half of those organisations surveyed have taken measures to develop digital preservation policies and to budget for it, while a majority have incorporated digital preservation into their organisational planning. Organisations predict that within a decade they will need to store large quantities of data in a wide range of formats from a variety of sources; three quarters of them are looking to invest in a solution within the next two years. However, the findings also point to varying degrees of readiness. Organisations with a digital preservation policy are significantly further advanced in their work to preserve digital collections for the long-term than others
Saving Second Life: Issues in Archiving a Complex, Multi-User Virtual World
Virtual environments, such as Second Life, have assumed an increasingly important role in popular culture, education and research. Unfortunately, we have almost no practical experience in how to preserve these highly dynamic, interactive information resources. This article reports on research by the National Digital Information Infrastructure for Preservation Program (NDIIPP)-funded Preserving Virtual Worlds project, which examines the issues that arise when attempting to archive regions from Second Life. Intellectual property and contractual issues can raise significant impediments to the creation of an archival information package for these environments, as can the technical design of the worlds themselves. We discuss the implication of these impediments for distributed models of preservation, such as NDIIPP
A Practice and Value Proposal for Doctoral Dissertation Data Curation
The preparation and publication of dissertations can be viewed as a subsystem of scholarly communication, and the treatment of data that support doctoral research can be mapped in a very controlled manner to the data curation lifecycle. Dissertation datasets represent “low-hanging fruit” for universities who are developing institutional data collections. The current workflow for processing electronic theses and dissertations (ETD) at a typical American university is presented, and a new practice is proposed that includes datasets in the process of formulating, awarding, and disseminating dissertations in a way that enables them to be linked and curated together. The value proposition and new roles for the university and its student-authors, faculty, graduate programs and librarians are explored