International Journal of Digital Curation
Not a member yet
605 research outputs found
Sort by
Pinning It Down: Towards a Practical Definition of ‘Research Data’ for Creative Arts Institutions
There is a widespread understanding among scientific researchers about what is meant by ‘research data’; however this does not readily translate into a creative context. As part of its engagement with the University of the Arts London (UAL) and via its support for the JISC Managing Research Data Programme, the Digital Curation Centre (DCC) and partners have worked towards an acceptable and practical definition of research data for creative arts institutions. This paper describes the activities carried out to help pin down such a definition, including a literature review, short and extended interviews with researchers, interactions with an academic arts research practitioner, and distillation of the results from a one-day workshop which took place in London in September 2012
Prenormative Research into Standard Messaging Formats for Engineering Materials Data
To qualify materials for specific applications comprehensive testing is necessary, and consequently the engineering materials community has developed an extensive collection of documentary testing standards to define test conditions, specimen configurations, and post processing and reporting procedures. Unfortunately, in the absence of corresponding data formats, test results are rarely conserved and their value diminishes as the material pedigree, test conditions, and results become disassociated. In an effort to address this issue, prenormative research has demonstrated the viability of deriving data formats from documentary testing standards and thus the possibility to realize a standards-based data infrastructure for the engineering materials community
Distributed Digital Preservation in the Cloud
The LOCKSS system is a leading technology in the field of Distributed Digital Preservation. Libraries run LOCKSS boxes to collect and preserve content published on the Web in PC servers with local disk storage. They form nodes in a network that continually audits their content and repairs any damage. Libraries wondered whether they could use cloud storage for their LOCKSS boxes instead of local disks. We review the possible configurations, evaluate their technical feasibility, assess their economic feasibility, report on an experiment in which we ran a production LOCKSS box in Amazon’s cloud service, and describe some simulations of future costs of cloud and local storage. We conclude that current cloud storage services are not cost-competitive with local hardware for long term storage, including for LOCKSS boxes
The Problematic Future of Research Data Management: Challenges, Opportunities and Emerging Patterns Identified by the DataRes Project
This paper describes findings and projections from a project that has examined emerging policies and practices in the United States regarding the long-term institutional management of research data. The DataRes project at the University of North Texas (UNT) studied institutional transitions taking place during 2011-2012 in response to new mandates from U.S. governmental funding agencies requiring research data management plans to be submitted with grant proposals. Additional synergistic findings from another UNT project, termed iCAMP, will also be reported briefly.This paper will build on these data analysis activities to discuss conclusions and prospects for likely developments within coming years based on the trends surfaced in this work. Several of these conclusions and prospects are surprising, representing both opportunities and troubling challenges, for not only the library profession but the academic research community as a whole
Evolving Persistent Archives and Digital Library Systems: Integrating iRods, Cheshire3 and Multivalent
This paper describes work undertaken by Data Intensive Cyber Environments Center (DICE) at the University of North Carolina at Chapel Hill and the University of Liverpool on the development of an integrated preservation environment, which has been presented at the National Coordination Office for Networking and Information Technology Research and Development (NITRD), at the National Science Foundation, and at the European Commission. The underlying technology is based on the integrated Rule-Oriented Data System (iRODS), which implements a policy-based approach to distributed data management. By differentiating between different phases of the data life cycle based upon the evolution of data management policies, the infrastructure can be tuned to support data publication, data sharing, data analysis and data preservation. It is possible to build generic data management infrastructure that can evolve to meet the management requirements of each user community, federal agency and academic research project. In order to manage the properties of the data collections, we have developed and integrated scalable digital library services that support the discovery of, and access to, material organized as a collection.The integrated preservation environment prototype implements specific technologies that are capable of managing a wide range of preservation requirements, from parsing of legacy document formats, to enforcement of preservation policies, to validation of trustworthiness assessment criteria. Each capability has been demonstrated and is instantiated in multiple instances, both in the United States as part of the DataNet Federation Consortium (DFC) and through multiple European projects, primarily the FP7 SHAMAN project
Creating a Research Data Management Service
This paper provides an overview of the elements required to create a sustainable research data management (RDM) service. The paper summarises key learning and lessons learnt from the University of Nottingham’s project to create an RDM service for researchers. Collective experiences and learning from three key areas are covered, including: data management requirements gathering and validation, RDM training, and the creation of an RDM website
The GESIS Data Archive for the Social Sciences: A Widely Recognised Data Archive on its Way
This paper describes initial experiences in evaluating an established data archive with a long-standing commitment to preservation and dissemination of social science research data against recently formulated standards for trustworthy digital archives. As stakeholders need to be sure that the data they produce, use or fund is treated according to common standards, the GESIS Data Archive decided to start a process of audit and certification within the European Framework of Certification and Audit, starting with the Data Seal of Approval (DSA). This paper gives an overview of workflows within the archive and illustrates some of the steps necessary to obtain the DSA as well as to optimize some of its services. Finally, a short appraisal of the method of the DSA is made
Editorial
Kevin Ashley, Chief Editor, introduces Volume 8, Issue 2 (2013) of the International Journal of Digital Curation
Data Management of Confidential Data
Social science researchers increasingly make use of data that is confidential because it contains linkages to the identities of people, corporations, etc. The value of this data lies in the ability to join the identifiable entities with external data, such as genome data, geospatial information, and the like. However, the confidentiality of this data is a barrier to its utility and curation, making it difficult to fulfil US federal data management mandates and interfering with basic scholarly practices, such as validation and reuse of existing results. We describe the complexity of the relationships among data that span a public and private divide. We then describe our work on the CED2AR prototype, a first step in providing researchers with a tool that spans this divide and makes it possible for them to search, access and cite such data
Understanding the ‘Intensive’ in ‘Data Intensive Research’: Data Flows in Next Generation Sequencing and Environmental Networked Sensors
Genomic and environmental sciences represent two poles of scientific data. In the first, highly parallel sequencing facilities generate large quantities of sequence data. In the latter, loosely networked remote and field sensors produce intermittent streams of different data types. Yet both genomic and environmental sciences are said to be moving to data intensive research. This paper explores and contrasts data flow in these two domains in order to better understand how data intensive research is being done. Our case studies are next generation sequencing for genomics and environmental networked sensors.Our objective was to enrich understanding of the ‘intensive’ processes and properties of data intensive research through a ‘sociology’ of data using methods that capture the relational properties of data flows. Our key methodological innovation was the staging of events for practitioners with different kinds of expertise in data intensive research to participate in the collective annotation of visual forms. Through such events we built a substantial digital data archive of our own that we then analysed in terms of three traits of data flow: durability, replicability and metrology.Our findings are that analysing data flow with respect to these three traits provides better insight into how doing data intensive research involves people, infrastructures, practices, things, knowledge and institutions. Collectively, these elements shape the topography of data and condition how it flows. We argue that although much attention is given to phenomena such as the scale, volume and speed of data in data intensive research, these are measures of what we call ‘extensive’ properties rather than intensive ones. Our thesis is that extensive changes, that is to say those that result in non-linear changes in metrics, can be seen to result from intensive changes that bring multiple, disparate flows into confluence.If extensive shifts in the modalities of data flow do indeed come from the alignment of disparate things, as we suggest, then we advocate the staging of workshops and other events with the purpose of developing the ‘missing’ metrics of data flow