1,720,972 research outputs found
Replication data and online supplement for: Underproduction: An Approach for Measuring Risk in Open Source Software
These materials were produced as part of:
Champion, Kaylea and Benjamin Mako Hill. (2021) "Underproduction: An approach for measuring risk in open source software.'' 28th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). Preprint: https://arxiv.org/abs/2103.00352. DOI: 10.1109/SANER50967.2021.00043
In this archive, you'll find:
inst_all_packages_full_results.tab Summary data on all packages as they appear in the paper. This is the place to look if you want to examine the underproduction factor associated with each package
inst_all_packages_full_results-DESCRIPTION.txt A description of the fields in the inst_all_packages_full_results.tab file.
R_Code.tar.gz Containing R code to reproduce figures and tables from fitted Bayesian hierarchical survival models:
dfPrep.R, used to create datasets_for_modeling.RData
models.R, a resource for model information
model_visualization.R, the core code for presenting fitted models and relationships
standalone_dsp.R, descriptive statistics
standalone_bayes.R, to produce tables for the paper
lib-00-utils.R, some utility functions
datasets_for_modeling.RData, the core dataset used for this analysis
Stan.tar.gz, a directory of STAN model output; on our supercomputing node these took multiple days to run and converge
Figures.tar.gz, a directory of figures from the paper
Raw_Data_Parsers.tar.gz, a directory of both the raw data and the parsers used to obtain the raw data. The dir contains a HowTo file if you would like to reproduce the scraping/cloning part of the project, however note that the original analysis included an rsync copy of the Debian bug database; if you conduct an analysis from scratch, the data you obtain will have changed since our rsync.
Appendix.tar.gz, containing figures and data associated with our appendix using an alternate measure of importance ("vote" which represents recent usage but omits packages where usage does not update atime; the paper used "inst")
appendix_with_vote.R, the code
appendix_figures, a directory of figures similar to those in the paper but produced for the appendix
vote_all_packages_full_results.csv -- summary data on all packages
vote_all_packages_full_results.csv.DESCRIPTION A description of the fields in the inst_all_packages_full_results.csv file.
For more information, please contact:
Kaylea Champion (she/her)
[email protected] | [email protected]
@kayleachampion
Abstract:
The widespread adoption of Free/Libre and Open Source Software (FLOSS) means that the ongoing maintenance of many widely used software components relies on the collaborative effort of volunteers who set their own priorities and choose their own tasks. We argue that this has created a new form of risk that we call `underproduction' which occurs when the supply of software engineering labor becomes out of alignment with the demand of people who rely on the software produced. We present a conceptual framework for identifying relative underproduction in software as well as a statistical method for applying our framework to a comprehensive dataset from the Debian GNU/Linux distribution that includes 21,902 source packages and the full history of 461,656 bugs. We draw on this application to present two experiments: (1) a demonstration of how our technique can be used to identify at-risk software packages in a large FLOSS repository and (2) a validation of these results using an alternate indicator of package risk. Our analysis demonstrates both the utility of our approach and reveals the existence of widespread underproduction in a range of widely-installed software components in Debian
Replication Data for: Taboo and Collaborative Knowledge Production: Evidence from Wikipedia
By definition, people are reticent or even unwilling to talk about taboo subjects. Because subjects like sexuality, health, and violence are taboo in most cultures, important information on each can be difficult to obtain. Are peer produced knowledge bases like Wikipedia a promising approach for providing people with information on taboo subjects? With its reliance on volunteers who might also be averse to taboo, can the peer production model be relied on to produce high-quality information on taboo subjects? In this paper, we seek to understand the role of taboo in volunteer-produced knowledge bases. We do so by developing a novel computational approach to identify taboo subjects and by using this method to identify a set of articles on taboo subjects in English Wikipedia. We find that articles on taboo subjects are more popular than non-taboo articles and that they are frequently subject to vandalism. Despite frequent attacks, we also find that taboo articles are higher quality. We hypothesize that societal attitudes will lead contributors to taboo subjects to seek to be less identifiable. Although our results are consistent with this proposal in several ways, we surprisingly find that contributors make themselves more identifiable in others
Replication Data for: Countering Underproduction of Information Public Goods
Information public goods are susceptible to a risk called “underproduction”: a mis-
alignment of the quality of the good and the demand for it. The practical con-
sequences vary by the type of good. When the information public good is digital
infrastructure, underproduced goods may be plagued by widescale security vulner-
abilities, and when the information public good is detailed factual information, un-
derproduced goods spread misinformation and ignorance. Yet contributors to in-
formation public goods are often volunteers who choose their own tasks, and are
not assigned to track consumer demand or to serve public needs. Although under-
production has been shown to be a common feature of peer-produced information
public goods, very little is known about how it can be counteracted. To understand
whether contributor experience and limits on their personal identifiability might
serve to counteract underproduction, we use a detailed longitudinal dataset from
English Wikipedia. First, we show that more experienced editors tend to contribute
to underproduced articles. Second, we show that this tendency to contribute in ways
that counter underproduction is still present although generally weaker among those
who contribute without an account. Third, we use a within-person analysis to show
that contributors tend to shift toward underproduced articles over time. These find-
ings illustrate the value of retaining peer production contributors, including those
contributing without accounts, as a means to counter underproduction
Python Libraries and the Shape of Collective Attention: Cybersecurity and the Rise of AI
The Python programming language is widely used across a range of fields and lines of business. How does interest in Python libraries change over time? Demand for goods like Python is difficult to measure, in part because Python is freely available as free/libre open source software (FLOSS). FLOSS is crucial to modern digital infrastructure such as the web and cloud, and numerous scientific endeavors including modeling and data mining. One part of sustaining digital infrastructure is better understanding how it is being used. In support of this, we consider two types of external events that may serve as demand triggers: internal events, in the form of announcements of security vulnerabilities, and external events, in the form of rising interest in machine learning and large language models in particular
Replication data for: Qualities of quality: A tertiary review of software quality measurement research
This archive contains supplemental materials for the following paper: "Defining the Qualities of Quality: A Systematic Review of Reviews of Software Quality Measurement"
In it, you'll find:
qualities_of_quality-dataset-human_readable.tsv — a human-friendly TSV containing fields extracted during the synthesis and review process
qualities_of_quality-dataset-r_friendly.tsv — an R-friendly version of the data extraction worksheet
data_ingest_and_clean.R,
descriptive_statistics.R, and
lib-00-utils.R — the R files used to calculate statistics and generate plots for the paper
qualities_of_quality-search_queries.pdf — a supplement describing our search queries
qualities_of_quality-article-20210729-DRAFT.pdf — a preprint of the article
</ul
Replication Materials for: Engineering Formality and Software Risk in Debian Python Packages
These materials were produced as part of:
Gaughan, Matthew, Champion, Kaylea, and & Hwang, Sohyeon. (2024) "Engineering Formality and Software Risk in Debian Python Packages." 31st IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER2024).
And includes data initially produced in:
Champion, Kaylea; Hill, Benjamin Mako, 2021, "Replication data and online supplement for: Underproduction: An Approach for Measuring Risk in Open Source Software", https://doi.org/10.7910/DVN/PUCD2P, Harvard Dataverse, V2
In this archive, you'll find:
inst_all_packages_full_results.tab Summary data for all Debian packages as they appear in Champion and Hill (2021).
mmt_data_final.csv Summary data for all Python-language Debian packages as they appear in the paper. This data set is novel, and includes package age in days, Github milestone usage, two different calculations of mean membership type (MMT), and package name.
calculatePower.R Contains R code to reproduce linear regression and power analysis methods as they appear in the paper.
For more information, please contact:
Matt Gaughan (he/him)
[email protected]
Abstract:
While Free/Libre and Open Source Software (FLOSS) is critical to global computing infrastructure, the maintenance of widely-adopted FLOSS packages is dependent on volunteer developers who select their own tasks. Risk of failure due to the misalignment of engineering supply and demand --- known as underproduction --- has led to code base decay and subsequent cybersecurity incidents such as the Heartbleed and Log4Shell vulnerabilities. FLOSS projects are self-organizing but can often expand into larger, more formal efforts. Although some prior work suggests that becoming a more formal organization decreases project risk, other work suggests that formalization may in fact increase the likelihood of project abandonment. We evaluate the relationship between underproduction and formality, focusing on formal structure, developer responsibility, and work processes management. We analyze 182 GNU/Linux packages made available via the Debian distribution and find that although more formal structures are associated with higher risk of underproduction, more elevated developer responsibility is associated with less underproduction while the relationship between formal work process management and underproduction is not statistically significant. Our analysis suggests that a FLOSS organization's transformation into a more formal structure may face unintended consequences which must be carefully managed. </p
Social and Technical Sources of Risk in Sustaining Digital Infrastructure
Thesis (Ph.D.)--University of Washington, 2024Significant risks to our shared digital infrastructure---communication systems, servers, and applications---can be identified by examining the social and technical conditions of the communities which produce that infrastructure. Exploration of these production communities reveals the deeply contingent processes of collective action that sustain them---processes that are innovative and powerful but sometimes fragile. As this shared body of digital infrastructure has grown, some crucial pieces have become neglected, leading to underproduction: the phenomenon of highly important, low-quality software packages. Underproduction is a form of what I will call a social production failure, and a substantial source of risk to digital infrastructure that today is used by billions of people. This dissertation is framed around a series of methodological and empirical projects. The first proposes a method for measuring underproduction risk in cross-section and demonstrates the application of that method to the Debian GNU/Linux community. Next, I examine the social and technical conditions of the Debian community and test hypotheses about how these conditions are associated with underproduction. I then develop a method to measure underproduction longitudinally, and apply this method to projects in Debian written using the Python programming language. I close by synthesizing these results with respect to my proposed theory of social production failures, and offer propositions and proposals for future work
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Thinking Deeply, Creating Richly: Learner Transformation Through Narrative
Narrative methods support transformative teaching and learning by accessing human cognitive strengths, including memory, reflection, and self-awareness. This paper explores the enduring and mindful use of narrative in education – as a method for transformative teaching and learning. A narrative is the intentional conversion of a group of events, participants, and details into a constructed reality that illustrates causes, characters, and results. Narrative development is a native human process by which we teach, learn, and remember. Narrative educational methods incorporate two key characteristics: integrative sense making, and shared connection building. Diverse disciplines – including biology, psychology, economics, literature, medicine, history, and education – have explored narrative as a foundational component of our human capacities, relationships, and achievements. Exploring the uses and misuses of narrative offers insight for teachers and learners of all ages. The paper closes with a discussion of the role this investigation is having in my personal and professional development
- …
