Association for the Advancement of Artificial Intelligence: AAAI Publications
Not a member yet
26155 research outputs found
Sort by
Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
Most web and digital trace data do not include information about an individual's nationality due to privacy concerns. The lack of data on nationality can create challenges for migration research. It can lead to a left-censoring issue since we are uncertain about the migrant's country of origin. Once we observe an emigration event, if we know the nationality, we can differentiate it from return migration. We propose methods to detect the nationality with the least available data, i.e., full names. We use the detected nationality in comparison with the country of academic origin, which is a common approach in studying the migration of researchers. We gathered 2.6 million unique name-nationality pairs from Wikipedia and categorized them into families of nationalities with three granularity levels to use as our training data. Using a character-based machine learning model, we achieved a weighted F1 score of 84% for the broadest- and 67%, for the most granular, country-level categorization. In our empirical study, we used the trained and tested model to assign nationality to 8+ million scholars' full names in Scopus data. Our results show that using the country of first publication as a proxy for nationality underestimates the size of return flows, especially for countries with a more diverse academic workforce, such as the USA, Australia, and Canada. We found that around 48% of emigration from the USA was return migration once we used the country of name origin in contrast to 33% based on academic origin. In the most recent period, 79% of scholars whose affiliation has consistently changed from the USA to China, and are considered emigrants, have Chinese names in contrast to 41% with a Chinese academic origin. Our proposed methods in addressing left-censoring issues are beneficial for other research that uses digital trace data to study migration
Polarized Online Discourse on Abortion: Frames and Hostile Expressions Among Liberals and Conservatives
Abortion has been one of the most divisive issues in the United States. Yet, missing is comprehensive longitudinal evidence on how political divides on abortion are reflected in public discourse over time, on a national scale, and in response to key events before and after the overturn of Roe v Wade. We analyze a corpus of over 3.5M tweets related to abortion over the span of one year (January 2022 to January 2023) from over 1.1M users. We estimate users' ideology and rely on state-of-the-art transformer-based classifiers to identify expressions of hostility and extract five prominent frames surrounding abortion. We use those data to examine (a) how prevalent were expressions of hostility (i.e., anger, toxic speech, insults, obscenities, and hate speech), (b) what frames liberals and conservatives used to articulate their positions on abortion, and (c) the prevalence of hostile expressions in liberals and conservative discussions of these frames. We show that liberals and conservatives largely mirrored each other's use of hostile expressions: as liberals used more hostile rhetoric, so did conservatives, especially in response to key events. In addition, the two groups used distinct frames and discussed them in vastly distinct contexts, suggesting that liberals and conservatives have differing perspectives on abortion. Lastly, frames favored by one side provoked hostile reactions from the other: liberals use more hostile expressions when addressing religion, fetal personhood, and exceptions to abortion bans, whereas conservatives use more hostile language when addressing bodily autonomy and women's health. This signals disrespect and derogation, which may further preclude understanding and exacerbate polarization
Measuring and Forecasting Conversation Incivility: the Role of Antisocial and Prosocial Behaviors
This paper focuses on the task of measuring and forecasting incivility in conversations following replies to hate speech. Identifying replies that steer conversations away from hatred and elicit civil follow-up conversations sheds light into effective strategies to engage with hate speech and proactively avoid further escalation. We propose new metrics that take into account various dimensions of antisocial and prosocial behaviors to measure the conversation incivility following replies to hate speech. Our best metric aligns with human perceptions better than prior work. Additionally, we present analyses on a) the language of antisocial and prosocial posts, b) the relationship between antisocial or prosocial posts and user interactions, and c) the language of replies to hate speech that elicit follow-up conversations with different incivility levels. We show that forecasting the incivility level of conversations following a reply to hate speech is a challenging task. We also present qualitative analyses to identify the most common errors made by our best model
Responsible RecSys by Design: Approximation Algorithms for Calibrated Recommendations with Sponsored Items
Calibrated Recommendation Systems (CRS) balance user preferences with constraints like diversity, fairness, and novelty to create inclusive recommendation lists. However, existing research often overlooks the mandatory inclusion of sponsored items, assuming unrestricted product selection. In practice, sponsored items, paid for by advertisers, must be included, which can conflict with CRS goals when advertisers' priorities misalign with system objectives. This paper addresses this gap by formulating CRS with sponsored items as a combinatorial optimization problem. We develop efficient approximation algorithms to generate the most calibrated recommendation lists while meeting sponsorship requirements
You Have the Floor: A Speaker-Aligned Corpus Derived from the Congressional Record
The United States Congressional Record serves as a comprehensive archive of legislative discourse, yet its sheer volume and unstructured format pose significant challenges for researchers interested in analyzing political language, speaker behavior, and ideological framing. This paper presents a new dataset that organizes Congressional speeches by individual speakers. Data was obtained using a mixture of the Congress.gov Application Programming Interface (API) and web-scraping techniques to retrieve the full text of the Congressional Record. After extracting roll-call votes and standardizing the transcripts to remove noisy artifacts and normalize formatting, each speaker is separated into individual files and annotated with metadata including name, political affiliation, years active, district or state represented, and professional social media accounts, if known. This enables fine-grained analysis of rhetorical patterns and linguistic strategies across different political groups as well as time periods. By making the dataset publicly available, we aim to support interdisciplinary research utilizing natural language processing (NLP)
The Suisse Romande Local News Dataset
This paper introduces a comprehensive dataset of news articles sourced from ESH Médias, a prominent local press agency in Romandy, the French-speaking region of Switzerland. The dataset encompasses all articles published on their digital platforms from January 2015 through June 2022. With over 130,000 articles written in French, this dataset offers a rich insight into local news from the French-speaking cantons of Switzerland. The articles cover a diverse range of topics and provide valuable material for Natural Language Processing and media studies. To respect privacy and legal considerations, journalists' names have been anonymized, and the dataset is made available for research purposes under a specific agreement with ESH Médias. The dataset adheres to the FAIR principles, and a detailed datasheet is provided to facilitate its use. The dataset is accessible via a DOI link
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
Online texts with toxic content are a clear threat to the users on social media in particular and society in general. Although many platforms have adopted various measures (e.g., machine learning-based hate-speech detection systems) to diminish their effect, toxic content writers have also attempted to evade such measures by using cleverly modified toxic words, so-called human-written text perturbations. Therefore, to help build automatic detection tools to recognize those perturbations, prior methods have developed sophisticated techniques to generate diverse adversarial samples. However, we note that these ``algorithm"-generated perturbations do not necessarily capture all the traits of ``human"-written perturbations. Therefore, in this paper, we introduce a novel, high-quality dataset of human-written perturbations, named as NoisyHate, that was created from real-life perturbations that are both written and verified by human-in-the-loop. We show that perturbations in NoisyHate have different characteristics than prior algorithm-generated toxic datasets show and thus can be particularly useful to help develop better toxic speech detection solutions. We also provide basic benchmark on the potential utilities of NoisyHate in perturbation normalization and understanding tasks. Both dataset and source code are publicly available
A General Paradigm for Fine-Tuning Large Language Models in Alzheimer’s Disease Diagnosis
Alzheimer's disease (AD), a complex neurodegenerative disorder, presents significant challenges for early and accurate diagnosis due to its multifactorial nature. This study introduces a novel approach to fine-tuning large language models (LLMs) for classifying AD-related dementia stages, using genetic and contextual demographic data. By harnessing the unique ability of LLMs to capture complex relationships in high-dimensional data, we developed a prompt structure that integrates genetic information, such as single nucleotide polymorphisms (SNPs), with patient-specific factors like age, sex, and clinical scores. Extensive experiments on the ADNI dataset demonstrate the superior performance of LLM-based methods. Our findings highlight the crucial role of high-quality prompts and carefully curated data in improving model accuracy. This research lays the groundwork for applying LLMs in precision medicine, providing a scalable and interpretable framework to address complex biomedical challenges, extending beyond AD
Team Context-Aware Collaborative AI Assistants
True collaboration or teamwork involves being aware of the team and its goals, adapting to the changing situations, communicating relevant information, identifying break-downs in common ground and repairing them, all to improve joint outcomes. At the heart of this type of collaborative agent are the capabilities needed to understand context and determine relevance. We introduce four key capabilities that enable our collaborative agents to be effective collaborators by understanding context and determining relevance
Revealing the Utilized Rank of Subspaces of Learning in Neural Networks
In this work, we study how well the learned weights of a neural network utilize the space available to them. This notion is related to capacity, but additionally incorporates the interaction of the network architecture with the dataset. Most learned weights appear to be full rank, and are therefore not amenable to low rank decomposition. This deceptively implies that the weights are utilizing the entire space available to them. We propose a simple data-driven transformation that projects the weights onto the subspace where the data and the weight interact. This preserves the functional mapping of the layer and reveals its low rank structure. In our findings, we conclude that most models utilize a fraction of the available space. For instance, for ViTB-16 and ViTL-16 trained on ImageNet, the mean layer utilization is 35% and 20% respectively. Our transformation results in reducing the parameters to 50% and 25% respectively, while resulting in less than 0.2% accuracy drop after fine-tuning. We also show that self-supervised pre-training drives this utilization up to 70%, justifying its suitability for downstream tasks