International Journal of Digital Curation
Not a member yet
605 research outputs found
Sort by
The world is all grown digital.... How shall a man persuade management what to do in such times?
Understanding and communicating the cost and value of digital curation activities has now been recognised by a number of projects and initiatives as a very important factor in ensuring the long-term survival of digital assets. A number of projects have developed costing models for digital preservation but there remains a major problem with information assets (digital or otherwise) in that their value is difficult to express in terms that are readily understood by all the stakeholders, especially those who might fund their preservation. This paper introduces a range of issues concerning information value and business models for sustained funding of digital preservation, with particular reference to the espida Project recently completed at the University of Glasgow. This project has developed a model of information value that builds on the Balanced Scorecard approach to business performance developed by Kaplan and Norton. This model casts information curation as an investment where current and ongoing expenditure is incurred in order to produce future returns, benefitting a range of stakeholders. In this formulation, value is seen as multi-facetted and, from the point of view of the individual or organisation funding the curation, explicitly related to the funder’s strategic goals. It also recognises that benefits may only accrue over the long term and that there is a risk that information that is preserved may fail to deliver any return. Examples discussed in the paper concern the establishment of an institutional repository and the establishment of an e-thesis service for an educational institution. It concludes that a deconstruction of benefits of this kind can be more quickly and fully understood even by stakeholders not necessarily expert in the curation field. This facilitates the production of a well-constructed case that clearly articulates information value and the benefit that accrues from its curation, which in turn allows senior management or other funders to make funding decisions based on understandable information: the basic premise of good practice in management. This is a commonly understood idea and one that the espida methodology helps fulfil
“The Naming of Cats”: Automated Genre Classification
This paper builds on the work presented at the ECDL 2006 in automated genre classification as a step toward automating metadata extraction from digital documents for ingest into digital repositories such as those run by archives, libraries and eprint services (Kim & Ross, 2006b). We have previously proposed dividing features of a document into five types (features for visual layout, language model features, stylometric features, features for semantic structure, and contextual features as an object linked to previously classified objects and other external sources) and have examined visual and language model features. The current paper compares results from testing classifiers based on image and stylometric features in a binary classification to show that certain genres have strong image features which enable effective separation of documents belonging to the genre from a large pool of other documents
User Priorities for Data: Results from SUPER
SUPER, a Study of User Priorities for e-infrastructure for Research, was a six-month effort funded by the UK e-Science Core Programme and JISC. Its aim was to inform investment in order to provide a usable, useful, and accessible e-infrastructure for all researchers and a coherent set of e-infrastructure services that would increase usage by at least a factor of ten by 2010. Through a series of unstructured face-to-face interviews with over 45 participants from 30 different projects, an online survey, together with a day-long workshop at NeSC, we have observed recurring issues relating to the provision of e-infrastructure. In this article we focus on the data-related issues identified during these interactions. We conclude with a prioritised list of future activities for research, development, and adoption in the data space
Editorial
Chris Rusbridge, Chief Editor, introduces the Summer 2008 issue of the International Journal of Digital Curation
Toward Distributed Infrastructures for Digital Preservation: The Roles of Collaboration and Trust
This paper first explores some of the reasons why collaboration is becoming increasingly important in supporting scientific data curation, digital preservation initiatives and institutional repository development. It then investigates the concepts of trust and control used in the organisation science literature and attempts to apply them to the work on trustworthy repositories being carried out by various international initiatives
Challenges and Directions in 3D and VR Data Curation: Findings from a Nominal Group Study
This study identifies challenges and promising directions in the curation of 3D data. 3D visualization shows great promise for a range of scholarly fields through interactive engagement with and analysis of spatially complex artifacts, spaces, and data. While the new affordability of emerging 3D capture technologies presents greater academic possibilities, academic libraries need more effective workflows, policies, standards, and practices to ensure that they can support the creation, discovery, access, preservation, and reproducibility of 3D data sets. This study uses nominal group technique with invited experts across several disciplines and sectors to identify common challenges in the creation and re-use of 3D data for the purpose of developing library strategy for supporting curation of 3D data. This article identifies staffing needs for 3D imaging; alignment with IT resources; the roll of archivists in addressing unique challenges posed by these datasets; the importance of data annotation, metadata, and transparency for research integrity and reproducibility; and features for storage, access, and management to facilitate re-use by researchers and educators. Participants identified three main challenges for supporting 3D data that align with the strengths of libraries: 1) development of crosswalks and aggregation tools for discipline-specific metadata models, data dictionaries for 3D research, and aggregation tools for expanding discovery; 2) development of an open source viewer that supports streaming and annotation on archival formats of 3D models and makes archival master files accessible, while also serving derivative files based on user requirements; and 3) widespread of adoption of better documentation and technical metadata for image capture and modeling processes in order to support replicability of research, reproducibility of models, and transparency of scientific process
Making Meaning of Historical Papua New Guinea Recordings: Collaborations of Speaker Communities and the Archive
PARADISEC’s PNG collections represent the great diversity in the regions and languages of PNG. In 2016 and 2017, in recognition of the value of PARADISEC’s collections, ANDS (the Australian National Data Service) provided funding for us to concentrate efforts on enhancing the metadata that describes our Papua New Guinea (PNG) collections, an effort designed to maximise the findability and useability of the language and music recordings preserved in the archive for both source communities and researchers. PARADISEC\u27s subsequent engagement with PNG language experts has led to collaborations with members of speaker communities who are part of the PNG diaspora in Australia. In this paper, we show that making historical recordings more findable, accessible and better described can result in meaningful interactions with and responses to the data in source communities. The effects of empowering speaker communities in their relationships to archives can be far reaching – even inverting, or disrupting the power relationships that have resulted from the colonial histories in which archives are embedded
Finding a Repository with the Help of Machine-Actionable DMPs: Opportunities and Challenges
Finding a suitable repository to deposit research data is a difficult task for researchers since the landscape consists of thousands of repositories and automated tool support is limited. Machine-actionable DMPs can improve the situation since they contain relevant context information in a structured and machine-friendly way and therefore enable automated support in repository recommendation.
This work describes the current practice of repository selection and the available support today. We outline the opportunities and challenges of using machine-actionable DMPs to improve repository recommendation. By linking the use case of repository recommendation to the ten principles for machine-actionable DMPs, we show how this vision can be realized. A filterable and searchable repository registry that provides rich metadata for each indexed repository record is a key element in the architecture described. At the example of repository registries we show that by mapping machine-actionable DMP content and data policy elements to their filter criteria and querying their APIs a ranked list of repositories can be suggested.
[This paper is a conference pre-print presented at IDCC 2020 after lightweight peer review.
Piloting a Community of Student Data Consultants that Supports and Enhances Research Data Services
Research ecosystems within university environments are continuously evolving and requiring more resources and domain specialists to assist with the data lifecycle. Typically, academic researchers and professionals are overcommitted, making it challenging to be up-to-date on recent developments in best practices of data management, curation, transformation, analysis, and visualization. Recently, research groups, university core centers, and Libraries are revitalizing these services to fill in the gaps to aid researchers in finding new tools and approaches to make their work more impactful, sustainable, and replicable. In this paper, we report on a student consultation program built within the University Libraries, that takes an innovative, student-centered approach to meeting the research data needs in a university environment while also providing students with experiential learning opportunities. This student program, DataBridge, trains students to work in multi-disciplinary teams and as student consultants to assist faculty, staff, and students with their real-world, data-intensive research challenges. Centering DataBridge in the Libraries allows students the unique opportunity to work across all disciplines, on problems and in domains that some students may not interact with during their college careers. To encourage students from multiple disciplines to participate, we developed a scaffolded curriculum that allows students from any discipline and skill level to quickly develop the essential data science skill sets and begin contributing their own unique perspectives and specializations to the research consultations. These students, mentored by Informatics faculty in the Libraries, provide research support that can ultimately impact the entire research process. Through our pilot phase, we have found that DataBridge enhances the utilization and openness of data created through research, extends the reach and impact of the work beyond the researcher’s specialized community, and creates a network of student “data champions” across the University who see the value in working with the Library. Here, we describe the evolution of the DataBridge program and outline its unique role in both training the data stewards of the future with regard to FAIR data practices, and in contributing significant value to research projects at Virginia Tech. Ultimately, this work highlights the need for innovative, strategic programs that encourage and enable real-world experience of data curation, data analysis, and data publication for current researchers, all while training the next generation of researchers in these best practices.
[This paper is a conference pre-print presented at IDCC 2020 after lightweight peer review.
Privacy Concerns in Qualitative Video Data Reuse
In this article, we examine how data producers’ and reusers’ privacy concerns shape their views about data sharing and reuse in the field of education, with an emphasis on video records of practice. We find that data producers and reusers were concerned about the risks that qualitative data, and video records of practice in particular, present to themselves, their colleagues, and the subjects represented in the data. Specifically, they emphasized risks relating to the privacy the subjects – teachers and students who appear in the videos. In response to these risks, data producers have engaged in a number of strategies to minimize risk and/or mitigate potential harm including: (1) education and training; (2) using informed consent to facilitate and/or restrict data sharing; and (3) limiting data capture/production. We discuss the implications that our findings have for digital repositories, and for efforts to facilitate the sharing and reuse of qualitative video data in education