Association for the Advancement of Artificial Intelligence: AAAI Publications
Not a member yet
26155 research outputs found
Sort by
Linguistic Landscape of Generative AI Perception: A Global Twitter Analysis Across 14 Languages
The advent of generative AI tools has had a profound impact on societies globally, transcending geographical boundaries. Understanding these tools' global reception and utilization is crucial for service providers and policymakers in shaping future policies. Therefore, to unravel the perceptions and engagements of individuals within diverse linguistic communities with regard to generative AI tools, we extensively analyzed over 6.8 million tweets in 14 different languages. Our findings reveal a global trend in the perception of generative AI, accompanied by language-specific nuances. While sentiments toward these tools vary significantly across languages, there is a prevalent positive inclination toward Image tools and a negative one toward Chat tools. Notably, the ban of ChatGPT in Italy led to a sentiment decline and initiated discussions across languages. Furthermore, we established a taxonomy for interactions with chatbots, creating a framework for social analysis underscoring variations in generative AI usage among linguistic communities. We find that the Chinese community predominantly employs chatbots as substitutes for search, while the Italian community tends to use chatbots for tasks such as problem-solving assistance and engaging in entertainment or creative tasks. Our research provides a robust foundation for further explorations of the social dynamics surrounding generative AI tools and offers invaluable insights for decision-makers in policy, technology, and education
Adopting Beliefs or Superficial Mimicry? Investigating Nuanced Ideological Manipulation of LLMs
Large Language Models (LLMs) have transformed natural language processing, but concerns have emerged about their susceptibility to ideological manipulation, particularly in politically sensitive areas. Previous research has largely focused on LLM biases through a binary Left vs. Right framework, often using explicit ideological prompts and fine-tuning with political question-answering datasets. In this work, we move beyond this binary approach to explore the extent to which LLMs can be influenced across a nuanced spectrum of political ideologies, from Progressive-Left to Conservative-Right. We introduce a novel multi-task dataset designed to reflect diverse ideological positions through tasks such as ideological question-answering, statement ranking, manifesto cloze completion, and Congress bill comprehension. By fine-tuning three LLMs—Phi-2, Mistral, and Llama-3—on this dataset, we evaluate their capacity to adopt and express these nuanced ideologies. Our findings indicate that fine-tuning significantly enhances nuanced ideological alignment, while explicit prompts provide only minor refinements. This highlights the models' susceptibility to subtle ideological manipulation, suggesting a need for more robust safeguards to mitigate these risks
Niche Dynamics in Complex Online Community Ecosystems
Online communities are important organizational forms where members socialize and share information. Curiously, different online communities often overlap considerably in topic and membership. Recent research has investigated competition and mutualism among overlapping online communities through the lens of organizational ecology; however, it has not accounted for how the nonlinear dynamics of online attention may lead to episodic competition and mutualism. Neither has it explored the origins of competition and mutualism in the processes by which online communities select or adapt to their niches. This paper presents a large-scale study of 8,806 Reddit communities belonging to 1,919 clusters of high user overlap over a 5-year period. The method uses nonlinear time series methods to infer bursty, often short-lived ecological dynamics. Results reveal that mutualism episodes are longer lived and slightly more frequent than competition episodes. Next, it tests whether online communities find their niches by specializing to avoid competition using panel regression models. It finds that competitive ecological interactions lead to decreasing topic and user overlaps; however, changes that decrease such niche overlaps do not lead to mutualism. The discussion considers future designs for online community ecosystem management
Characterizing Knowledge Manipulation in a Russian Wikipedia Fork
Wikipedia is powered by MediaWiki, a free and open-source software that is also the infrastructure for many other wiki-based online encyclopedias. These include the recently launched website Ruwiki, which has copied and modified the original Russian Wikipedia content to conform to Russian law. To identify practices and narratives that could be associated with different forms of knowledge manipulation, this article presents an in-depth analysis of this Russian Wikipedia fork. We propose a methodology to characterize the main changes with respect to the original version. The foundation of this study is a comprehensive comparative analysis of more than 1.9M articles from Russian Wikipedia and its fork. Using meta-information and geographical, temporal, categorical, and textual features, we explore the changes made by Ruwiki editors. Furthermore, we present a classification of the main topics of knowledge manipulation in this fork, including a numerical estimation of their scope. This research not only sheds light on significant changes within Ruwiki, but also provides a methodology that could be applied to analyze other Wikipedia forks and similar collaborative projects
Detecting Harassment and Defamation in Cyberbullying with Emotion-Adaptive Training
Existing research on detecting cyberbullying incidents on social media has primarily concentrated on harassment and is typically approached as a binary classification task. However, cyberbullying encompasses various forms, such as denigration and harassment, which celebrities frequently face. Furthermore, suitable training data for these diverse forms of cyberbullying remains scarce. In this study, we first develop a celebrity cyberbullying dataset that encompasses two distinct types of incidents: harassment and defamation. We investigate various types of transformer-based models, namely masked (RoBERTa, Bert and DistilBert), replacing (Electra), autoregressive (XLnet), masked&permuted (Mp-net), text-text (T5) and large language models (Llama2 and Llama3) under low source settings. We find that they perform competitively on explicit harassment binary detection, however, their performance is substantially lower on harassment and denigration multi-classification tasks. Therefore,
we propose an emotion-adaptive training framework (EAT) that helps transfer knowledge from the domain of emotion detection to the domain of cyberbullying detection to help detect indirect cyberbullying events. EAT consistently improves
the average macro F1, precision and recall by 20% in cyberbullying detection tasks across nine transformer-based models under low-resource settings. Our claims are supported by intuitive theoretical insights and extensive experiments
Examining the Makeup of Media Trigger Warnings Online
In today’s digital landscape, the prevalence of sensitive online content has made trigger warnings essential. These warnings inform viewers that the content they are about to see contains sensitive artifacts (e.g. violence). This paper studies the use of trigger warnings, exploiting data from two major platforms: Does the Dog Die, a crowdsourcing trigger warnings platform, and IMDb, a media database. We first study how different media types (e.g. films, video games, and TV shows) are labeled with varying trigger warnings and the different co-occurrence patterns among different trigger warnings. We also discover controversy surrounding certain trigger warnings, with inconsistent opinions stated by different people. We further show that different jurisdictions (e.g. USA vs. UK) assign different content ratings (e.g. R-18) for the same media, even when the same trigger warnings are present. Finally, we develop automatic detectors to identify trigger warnings from IMDb text. We achieve F1 scores exceeding 0.7 for all 10 selected trigger warnings
ArDia: Improving Arabic Dialectal Language Classification Using a Novel Dataset
Despite Arabic being one of the most widely spoken languages, there is a scarcity of available dialectal Arabic data. In this paper, we address this challenge by proposing a novel approach to data collection through the main use of video captions from TikTok, and other resources such as dictionaries and articles, resulting in the creation of the ArDia dataset. To the best of our knowledge, the ArDia dataset is the largest labeled dialectal Arabic dataset, containing over 900,000 examples, each labeled with its respective dialect. We further leverage this dataset to pretrain transformer-based models, ArDiaBERT and ArDiaGPT. Due to a lack of research on the Arabic models, we present a comprehensive study of Arabic dialect identification using the ArDia dataset on the dialect identification task
Survival Analysis for Cancers of the Brain, CNS and Bone using Retrieval Augmented Generation on the SEER Database
Mortality estimation remains a key issue in cancers affecting the Brain, Central Nervous System (CNS), and Bone, among others. The recent integration of LLM-based reasoning into tools that aid cancer prognosis has been particularly encouraging. This prompts us to examine further their stated efficacy and devise workarounds to reduce hallucinations using retrieval augmented generation. We study the clinical, pathological and demographic logs of patients recorded in the National Institutes of Health (NIH) Surveillance, Epidemiology, and End Results (SEER) database and develop an integrated methodology that is user-friendly and responds to n-shot queries with or without context. We first build a set of custom SEER embeddings using DistilBERT, which we use to test tree-based models in answering 'yes/no' type 5-year survivability questions given patient profiles. We extend the limited binary response capability of the prior models by using TabLLM, HyDE-RAG, and Step-Back RAG on the BCNS cancer data and extend them to Bone Cancer data from SEER using GraphRAG, as the attributes are similar. The conversation-friendly models are able to take different context lengths and types into account and provide reasoning about their responses. We successfully show that the extensive patient records in the SEER database can be utilized to develop a powerful conversational agent that is not only able to classify mortality outcomes but also reason about the response by leveraging latent inter-relationships among the unique clinical variables
Bidirectional Feedback-Based Personalization of Learning using Multi-tier AI: A Real-World Assessment of its Efficacy in Classrooms
Jill Watson is an example of an intelligent conversational AI Teaching Assistant that has been deployed across 24 class sections in different institutions, with 1102 unique student participants and over 17000 questions from Fall 2023 to present. Jill Watson’s RAG-based architecture built around OpenAI ChatGPT and user study results addresses some of the concerns related to domain knowledge, deployment, and data collection in online classrooms. In this work, a 3-tiered framework for personalization in online education using AI tools to enable a human-AI personalization loop, grounded in real-world human feedback is proposed
AI-Mediated Dispute Resolution
We examine the effectiveness of large language model (LLM) mediations in the under-studied dispute resolution domain. We first used a new corpus of dispute resolutions, KODIS, to investigate if LLMs can correctly identify whether to intervene. We find evidence that GPT as a mediator picks up on salient aspects of a dispute, such as Frustration and whether the disputants ultimately come to a resolution or stall at an impasse --- intervening significantly more so in cases of high frustration and impasse. Afterward, we ran a user study to compare GPT mediations against those of novice human mediators. We find participants agreed GPT's mediations were more likely to lead to resolution; were better positioned in the dialog; had better justification than human-crafted ones; and, on a forced choice, were generally more effective than novice human mediations