9607 research outputs found
Sort by
Enabling Embodied Music-Making for Non-Musicians
We present a Research through Design exploration of the potential for using tangible and embodied interactions to enable active music experiences - musicking - for non-musicians. We present the Tubularium prototype, which aims to facilitate music-making to non-musicians by not requiring any initial skill while still eliciting agency and overall, providing a meaningful experience. We present the design of the prototype and the features implemented and reflect on insights from a public event in which the prototype was trialed
Multimodal Extraction and Recognition of Arabic Implicit Discourse Relations
Most research on implicit discourse relation identification has focused on written language, however, it is also crucial to understand these relations in spoken discourse. We introduce a novel method for implicit discourse relation identification across both text and speech, that allows us to extract examples of semantically equivalent pairs of implicit and explicit discourse markers, based on aligning speech+transcripts with subtitles in another language variant. We apply our method to Egyptian Arabic, resulting in a novel high-quality dataset of spoken implicit discourse relations. We present a comprehensive approach to modeling implicit discourse relation classification using audio and text data with a range of different models. We find that text-based models outperform audio-based models, but combining text and audio features can lead to enhanced performance
Formant-Based Vowel Categorization for Cross-Lingual Phone Recognition
Multilingual phone recognition models can learn language-independent pronunciation patterns from large volumes of spoken data and recognize them across languages. This potential can be harnessed to improve speech technologies for underresourced languages. However, these models are typically trained on phonological representations of speech sounds, which do not necessarily reflect the phonetic realization of speech. A mismatch between a phonological symbol and its phonetic realizations can lead to phone confusions and reduce performance. This work introduces formant-based vowel categorization aimed at improving cross-lingual vowel recognition by uncovering a vowel's phonetic quality from its formant frequencies, and reorganizing the vowel categories in a multilingual speech corpus to increase their consistency across languages. The work investigates vowel categories obtained from a trilingual multi-dialect speech corpus of Danish, Norwegian, and Swedish using three categorization techniques. Cross-lingual phone recognition experiments reveal that uniting vowel categories of different languages into a set of shared formant-based categories improves cross-lingual recognition of the shared vowels, but also interferes with recognition of vowels not present in one or more training languages. Cross-lingual evaluation on regional dialects provides inconclusive results. Nevertheless, improved recognition of individual vowels can translate to improvements in overall phone recognition on languages unseen during training
A Corrosive Decline
This piece is written in the wake of damaging projects of extractive capitalism. As the Trump administration pulls back government support for clean energy, energy policy and market forces are past the tipping point of change. Trump brings uncertainty to the structural reorganization of the global economy, yet he will not reverse it. Writing from Denmark, David Struthers looks to the Indigenous analyses of ongoing colonialism to make an argument against a corrosive authoritarian encroachment
Iterative Structured Knowledge Distillation: Optimizing Language Models Through Layer-by-Layer Distillation
Traditional language model compression techniques, like knowledge distillation, require a fixed architecture, limiting flexibility, while structured pruning methods often fail to preserve performance. This paper introduces Iterative Structured Knowledge Distillation (ISKD), which integrates knowledge distillation and structured pruning by progressively replacing transformer blocks with smaller, efficient versions during training. This study validates ISKD on two transformer-based language models: GPT-2 and Phi-1. ISKD outperforms L1 pruning and achieves similar performance to knowledge distillation while offering greater flexibility. ISKD reduces model parameters - 30.68% for GPT-2 and 30.16% for Phi-1 - while maintaining at least four-fifths of performance on both language modeling and commonsense reasoning tasks. These findings suggest that this method offers a promising balance between model efficiency and accuracy
Subword symmetry in natural languages
Symmetric patterns are found in the orderly arrangements of natural structures, from proteins to the symmetry in animals’ bodies. Symmetric structures are more stable and easier to describe and compress, which is why they may have been preferred as building blocks in natural selection. The idea that natural languages undergo an evolutionary process akin to the evolution of species has been pervasive in the study of language. This process might result in symmetric patterns as in other natural structures, but the notion of symmetry is rarely associated with the study of natural language. In this study, we look for symmetric patterns in text data, considering the length of subword units under a range of possible subword analyses. We study the length of subword units in 32 languages and discover that the splits of long words tend to be symmetric regardless of the segmentation method and that some automatic methods give symmetric splits at all word lengths. These results include natural language in the set of phenomena that can be described in terms of symmetry, opening a new research avenue for the empirical study of text data as a structure comparable to various other structures in the natural world
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
Socioeconomic status (SES) fundamentally influences how people interact with each other and, more recently, with digital technologies like large language models (LLMs). While previous research has highlighted the interaction between SES and language technology, it was limited by reliance on proxy metrics and synthetic data. We survey 1,000 individuals from ‘diverse socioeconomic backgrounds’ about their use of language technologies and generative AI, and collect 6,482 prompts from their previous interactions with LLMs. We find systematic differences across SES groups in language technology usage (i.e., frequency, performed tasks), interaction styles, and topics. Higher SES entail a higher level of abstraction, convey requests more concisely, and topics like ‘inclusivity’ and ‘travel’. Lower SES correlates with higher anthropomorphization of LLMs (using ”hello” and ”thank you”) and more concrete language. Our findings suggest that while generative language technologies are becoming more accessible to everyone, socioeconomic linguistic differences still stratify their use to create a digital divide. These differences underscore the importance of considering SES in developing language technologies to accommodate varying linguistic needs rooted in socioeconomic factors and limit the AI Gap across SES groups
Summon a demon and bind it: A grounded theory of LLM red teaming
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a formal qualitative methodology, we interviewed dozens of practitioners from a broad range of backgrounds, all contributors to this novel work of attempting to cause LLMs to fail. We focused on the research questions of defining LLM red teaming, uncovering the motivations and goals for performing the activity, and characterizing the strategies people use when attacking LLMs. Based on the data, LLM red teaming is defined as a limit-seeking, non-malicious, manual activity, which depends highly on a team-effort and an alchemist mindset. It is highly intrinsically motivated by curiosity, fun, and to some degrees by concerns for various harms of deploying LLMs. We identify a taxonomy of 12 strategies and 35 different techniques of attacking LLMs. These findings are presented as a comprehensive grounded theory of how and why people attack large language models: LLM red teaming
Lex Loot Boxes - The Regulation of Gambling-like Products in Video Games
Loot boxes are gambling-like products inside video games that players can purchase with real-world money to obtain random rewards. Most of the time, the player receives an undesirable reward and ‘loses’. Occasionally, the player receives a highly sought-after prize. Such randomisation means players usually need to make repeated purchases and spend a substantial sum of money if they wish to obtain certain specific rewards. Many have argued that loot boxes are similar to traditional gambling, both structurally and psychologically. A link has also been found between spending money on loot boxes and experiencing harms from problem gambling, among other risk factors. The scientific literature requires time to develop before it can provide more definitive answers as to the potential harms of loot boxes. However, these products are of concern to players, parents, regulators, and policymakers given their resemblance to traditional gambling, coupled with their availability to young children and adults alike – without proper regulation (if any). This has led several countries to consider and adopt regulation to proactively address the problem. The video game industry has also adopted certain self-regulatory rules of limited efficacy in response to appease the public and perhaps forestall stricter and more effective regulation. Grounded in psychological and sociological findings on the potential harms of loot boxes, this thesis explores potential regulatory approaches in terms of what various countries have done or proposed to do. In accordance with the principles of open science, the thesis empirically and transparently evaluates how well various measures in different countries have been implemented in practice, i.e., whether companies have complied with them. Taken together, this body of work demonstrates that compliance with loot box regulation has generally been poor across the world. The regulation of loot boxes is unsurprisingly difficult, given the massive volume of content that must be monitored. Companies should seek to better understand the rules and comply more effectively while regulators should better inform companies and more actively enforce the rules, including initially using less formal and cheaper methods of enforcement. Given the lack of government funding for regulators and how this is unlikely to change in the foreseeable future, the most effective method of enhancing and (hopefully) ensuring compliance appears to be academic advocacy research that directly impacts upon policymaking and implementation, such as this thesis