1,721,076 research outputs found
Tweet Acts: A Speech Act Classifier for Twitter
Speech acts are a way to conceptualize speech as action. This holds true for communication on any platform, including social media platforms such as Twitter. In this paper, we explored speech act recognition on Twitter by treating it as a multi-class classification problem. We created a taxonomy of six speech acts for Twitter and proposed a set of semantic and syntactic features. We trained and tested a logistic regression classifier using a data set of manually labelled tweets. Our method achieved a state-of-the-art performance with an average F1 score of more than 0.70. We also explored classifiers with three different granularities (Twitter-wide, type-specific and topic-specific) in order to find the right balance between generalization and overfitting for our task
New horizons in the study of child language acquisition
URL to paper on conference site.Naturalistic longitudinal recordings of child development promise to reveal fresh perspectives on fundamental questions of language acquisition. In a pilot effort, we have recorded 230,000 hours of audio-video recordings spanning the first three years of one child's life at home. To study a corpus of this scale and richness, current methods of developmental cognitive science are inadequate. We are developing new methods for data analysis and interpretation that combine pattern recognition algorithms with interactive user interfaces and data visualization. Preliminary speech analysis reveals surprising levels of linguistic fine-tuning by caregivers that may provide crucial support for word learning. Ongoing analyses of the corpus aim to model detailed aspects of the child's language development as a function of learning mechanisms combined with lifetime experience. Plans to collect similar corpora from more children based on a transportable recording system are underway.National Science Foundation (U.S.)MIT Center for Future BankingMassachusetts Institute of Technology. Media LaboratoryUnited States. Office of Naval ResearchUnited States. Dept. of Defens
A Semi-Automatic Method for Efficient Detection of Stories on Social Media
Twitter has become one of the main sources of news for many people. As real-world events and emergencies unfold,Twitter is abuzz with hundreds of thousands of stories about the events. Some of these stories are harmless, while others could potentially be life saving or sources of malicious rumors. Thus, it is critically important to be able to efficiently track stories that spread on Twitter during these events. In this paper, we present a novel semi-automatic tool that enables users to efficiently identify and track stories about real-world events on Twitter. We ran a user study with 25 participants, demonstrating that compared to more conventional methods, our tool can increase the speed and the accuracy with which users can track stories about real-world events
Behavior compilation for AI in games
In order to cooperate effectively with human players, characters need to infer the tasks players are pursuing and select contextually appropriate responses. This process of parsing a serial input stream of observations to infer a hierarchical task structure is much like the process of compiling source code. We draw an analogy between compiling source code and compiling behavior, and propose modeling the cognitive system of a character as a compiler, which tokenizes observations and infers a hierarchical task structure. An evaluation comparing automatically compiled behavior to human annotation demonstrates the potential for this approach to enable AI characters to understand the behavior and infer the tasks of human partners.Singapore-MIT GAMBIT Game La
Mapping Twitter Conversation Landscapes
While the most ambitious polls are based on standardized interviews
with a few thousand people, millions are tweeting freely and publicly in their own voices about issues they care about. This data offers a vibrant 24/7 snapshot of people’s response to various events and topics. The sheer scale of the data on Twitter allows us to measure in aggregate how the
various issues are rising and falling in prominence over time. However, the volume of the data also means that an intelligent tool is required to allow the users to make sense of the data. To this end, we built a novel, interactive web-based tool for mapping the conversation landscapes on Twitter. Our system utilizes recent advances in natural language processing and deep neural networks that are robust with respect to the noisy and unconventional nature of tweets, in conjunction with a scalable clustering algorithm an interactive visualization engine to allow users to tap the mine of information that is Twitter. We ran a user study with 40 participants using tweets about the 2016 US presidential election and the summer 2016 Orlando shooting, demonstrating that compared to more conventional methods, our tool can increase the speed and the accuracy with which users can identify and make sense of the various conversation topics on Twitter
Automatic Detection and Categorization of Election-Related Tweets
With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public nature of Twitter, political scientists now potentially have the means to analyse and understand the narratives that organically form, spread and decline among the public in a political campaign.However, the volume and diversity of the conversation on Twitter, combined with its noisy and idiosyncratic nature, make this a hard task. Thus, advanced data mining and language processing techniques are required to process and analyse the data. In this paper, we present and evaluate a technical framework, based on recent advances in deep neural networks, for identifying and analysing election-related conversation on Twitter on a continuous, longitudinal basis. Our models can detect election-related tweets with an F-score of 0.92 and can categorize these tweets into 22 topics with an F-score of 0.90
Semi-automated dialogue act classification for situated social agents in games
As a step toward simulating dynamic dialogue between agents and humans in virtual environments, we describe learning a model of social behavior composed of interleaved utterances and physical actions. In our model, utterances are abstracted as {speech act, propositional content, referent} triples. After training a classifier on 100 gameplay logs from The Restaurant Game annotated with dialogue act triples, we have automatically classified utterances in an additional 5,000 logs. A quantitative evaluation of statistical models learned from the gameplay logs demonstrates that semi-automatically classified dialogue acts yield significantly more predictive power than automatically clustered utterances, and serve as a better common currency for modeling interleaved actions and utterances
Semantic context effects on color categorization
A number of recent theories of semantic representation propose two-way interaction of semantic and perceptual information. These theories are supported by a growing body of experiments that show widespread interactivity of semantic and perceptual levels of processing. In the current experiment, participants classified ambiguous colors into one of two color categories. The ambiguous colors were presented either as a color patch (no semantic context), as an icon representing an object that is strongly associated with one of the color options, or as a word referring to such an object. Although the iconic and lexical contexts were incidental and irrelevant to the color categorization task, participants' responses were consistently biased toward the context color. These results extend previous findings by showing that lexical contexts, as well as iconic contexts influence color categorization.National Science Foundation (U.S.) (grant F32HD052364)National Science Foundation (U.S.) (Graduate Research Fellowship
Understanding speech in interactive narratives with crowd sourced data
Speech recognition failures and limited vocabulary coverage pose challenges for speech interactions with characters in games. We describe an end-to-end system for automating characters from a large corpus of recorded human game logs, and demonstrate that inferring utterance meaning through a combination of plan recognition and surface texts similarity compensates for recognition and understanding failures significantly better than relying on surface similarity alone.Singapore-MIT GAMBIT Game La
Fast transcription of unstructured audio recordings
URL to conference session list. Title is under heading: Wed-Ses1-P1:
Phonetics, Phonology, cross-language comparisons, pathologyWe introduce a new method for human-machine collaborative speech transcription that is significantly faster than existing transcription methods. In this approach, automatic audio processing algorithms are used to robustly detect speech in audio recordings and split speech into short, easy to transcribe segments. Sequences of speech segments are loaded into a transcription interface that enables a human transcriber to simply listen and type, obviating the need for manually finding and segmenting speech or explicitly controlling audio playback. As a result, playback stays synchronized to the transcriber's speed of transcription. In evaluations using naturalistic audio recordings made in everyday home situations, the new method is up to 6 times faster than other popular transcription tools while preserving transcription quality
- …
