1,721,101 research outputs found
Arabic dialect identification under scrutiny:Limitations of single-label classification
Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared tasks have been running since 2018. However, ADI systems are reported to fail in distinguishing between the micro-dialects of Arabic. We argue that the currently adopted framing of the ADI task as a single-label classification problem is one of the main reasons for that. We highlight the limitation of the incompleteness of the Dialect labels and demonstrate how it impacts the evaluation of ADI systems. A manual error analysis for the predictions of an ADI, performed by 7 native speakers of different Arabic dialects, revealed that ≈ 67% of the validated errors are not true errors. Consequently, we propose framing ADI as a multi-label classification task and give recommendations for designing new ADI datasets
Exploring Author Context for Detecting Intended vs Perceived Sarcasm
We investigate the impact of using author context on textual sarcasm detection. We define author context as the embedded representation of their historical posts on Twitter and suggest neural models that extract these representations. We experiment with two tweet datasets, one labelled manually for sarcasm, and the other via tag-based distant supervision. We achieve state-of-the-art performance on the second dataset, but not on the one labelled manually, indicating a difference between intended sarcasm, captured by distant supervision, and perceived sarcasm, captured by manual labelling.<br/
Exploring Author Context for Detecting Intended vs Perceived Sarcasm
We investigate the impact of using author context on textual sarcasm detection. We define author context as the embedded representation of their historical posts on Twitter and suggest neural models that extract these representations. We experiment with two tweet datasets, one labelled manually for sarcasm, and the other via tag-based distant supervision. We achieve state-of-the-art performance on the second dataset, but not on the one labelled manually, indicating a difference between intended sarcasm, captured by distant supervision, and perceived sarcasm, captured by manual labelling.<br/
Expression and perception of identity through skin-toned emoji
The introduction of emoji skin tone modifiers to the Unicode Standard in 2015 was met with considerable debate on the extent to which these emoji would be used, who would actually use them, and what they would actually be used for. I evaluate such claims against large datasets drawn from social media and find evidence that people generally produce skin-toned emoji which align with their real-world identity. I also identify particular variations in emoji production based on real-life skin tone and geographical location, as well as the context in which emoji are used. I test experimentally the extent to which people perceive identity from the emoji that others produce, finding that these emoji are strongly considered to represent specific identities for both authors and readers of social media. Even default yellow emoji without a skin tone are associated with a particular identity. The ability of emoji to index identity for both author and reader is a property found in language, where one consequence of this is a difference in attitudes, responses or behaviours towards perceived identities. Whether the indexicality of emoji can similarly affect behavioural outcomes is tested experimentally, where no such effect is found
Computational sarcasm detection and understanding in online communication
The presence of sarcasm in online communication has motivated an increasing number of computational investigations of sarcasm across the scientific community. In this thesis, we build upon these investigations. Pointing out their limitations, we bring four contributions that span two research directions: sarcasm detection and sarcasm understanding.
Sarcasm detection is the task of building computational models optimised for recognising sarcasm in a given text.
These models are often built in a supervised learning paradigm, relying on datasets of texts labelled for sarcasm.
We bring two contributions in this direction.
First, we question the effectiveness of previous methods used to label texts for sarcasm. We argue that the labels they produce might not coincide with the sarcastic intention of the authors of the texts that they are labelling.
In response, we suggest a new method, and we use it to build iSarcasm, a novel dataset of sarcastic and non-sarcastic tweets.
We show that previous models achieve considerably lower performance on iSarcasm than on previous datasets, while human annotators achieve a considerably higher performance, compared to models, pointing out the need for more effective models.
Therefore, as a second contribution, we organise a competition that invites the community to create such models.
Sarcasm understanding is the task of explicating the phenomena that are subsumed under the umbrella of sarcasm through computational investigation.
We bring two contributions in this direction.
First, we conduct an alaysis into the socio-demographic ecology of sarcastic exchanges between human interlocutors. We find that the effectiveness of such exchanges is influenced by the socio-demographic similarity between the interlocutors, with factors such as English language nativeness, age, and gender, being particualry influential. We suggest that future social analysis tools should account for these factors.
Second, we challenge the motivation of a recent endeavour of the community; mainly, that of augmenting dialogue systems with the ability to generate sarcastic responses. Through a series of social experiments, we provide guidelines for dialogue systems concerning the appropriateness of generating sarcastic responses, and the formulation of such responses.
Through our work, we aim to encourage the community to consider computational investigations of sarcasm interdisciplinarily, at the intersection of natural language processing and computational social science
Analysing privacy in online social media
People share a wide variety of information on social media, including personal and sensitive information, without understanding the size of their audience which may cause privacy complications. The networked nature of the platforms further exacerbates these complications where the information can be shared without the information owner's control. People also struggle to achieve their intended audience using the privacy settings provided by the platforms. In this thesis, I analyse potential privacy violations caused by social media users and their networks, as well as the usage and understanding of privacy settings. I focus on Twitter which has rather simplistic privacy settings with binary states.
The first part of my studies includes investigating personal information disclosures by networks using congratulatory messages. I analyse these messages and detect various types of life events including relationships, illness, familial matters, and birthdays. I show that public replies are enough to infer the content of the original message, even if the event subject hides or deletes the message. I further focus on birthdays which is one of the most popular life events and the potential date of birth disclosure has security implications besides the privacy ones. I show that over 1K users have their date of birth exposed daily, where 10% of these users have protected their tweets. I also show that users react positively to these congratulatory messages even though these posts potentially disclose personal and sensitive information.
In the second part of my thesis, I focus on privacy settings on Twitter. I quantify the usage patterns of privacy settings and investigate the reasons for changing these settings between public and protected by conducting a mixed-method study. I show that there is a set of users who frequently utilize the privacy settings provided by the platform. I also show that users turn protected to share personal content and regulate boundaries while they turn public to interact with others in ways prevented by being protected.
In the last stage of the thesis, I investigate the user understanding of information and tweet visibility of different account types by conducting a user survey. I show that the users are aware of the visibility of their profile information and individual tweets. However, the visibility of followed topics, lists, and interactions with protected accounts is confusing. Less than a third of the survey participants were aware that a reply by a public account to a protected account's tweet would be publicly visible. Surprisingly, having a protected account did not result in a better understanding of the information or tweet visibility.
Actual functionalities and the user understanding of them should align so that users can take the right actions for desired levels of privacy protection in online social networks. I show that even with simplistic privacy settings, users have difficulty understanding the reach of their posts. Implications of interactions between users need to be clearly relayed. I give design suggestions to increase this awareness and for users to have better tools to manage their boundaries. I conclude the thesis by giving general implications around the studies conducted and possible future directions
Spiritual polarisation on social media: the case of Arab atheists on Twitter
Social media platforms provide an unprecedented method of communication, and they are considered an integral part of people's lifestyles. Also, these platforms facilitate forming communities, groups and networks. Hence, it attracted researchers to study people's interactions and analyse the enormous human-generated data. In this thesis, I focus on studying the online Arab communities as a case study of online communities to understand online spiritual-based groups and the polarisation among them. This work combines multi-disciplinary approaches of natural language processing, information retrieval, data science and social and technological networks to understand better the online social behaviour of Arabs with different religious beliefs. I explore the discussion among Arab Twitter users from religious and atheistic groups. I identify four types of Twitter users based on how they describe themselves: Atheistic, Theistic, Tanweeri (reformers), and Rationalists. This study shows that Arabs from different religious spectrums get involved in online discussions on local and regional topics.
I collected two datasets from Twitter for users who discussed religions and atheism, in which I considered about 434 accounts in the first dataset and 2,673 accounts in the second one. The analysis shows that, whatever their attitude towards religions, Arab Twitter users tend to use their accounts to promote their beliefs and to show their stances towards others. I showed that the data that was generated by these four groups illustrate the rich socio-cultural context in which discussions among believers, non-believers and religious reformers unfold. I showed that there is a clear online polarisation between atheists and theists, while Rationalist and Tanweeri accounts are spread among and between the two polarised groups. Arab atheists are separated into two groups in terms of engagement based on the accounts they prefer to interact-with.
I found that Arab atheists and theists mention and reply-to users from any religious groups and vice versa, but they tend to retweet and follow accounts from their own group. The findings of this thesis provide insights for researchers to understand the case study of Arab online communities and the religious and non-religious online polarisation. Also, it shows the implications for the studies of spiritual discourse on social media and provides a better cross-cultural understanding of relevant aspects
Opinion summarization of multiple reviews: data synthesis and modeling
The proliferation of online reviews has accelerated research on opinion mining, where
the ultimate goal is to glean information from reviews which help users make decisions
more efficiently. While opinion mining has assumed several facets in the literature
(e.g., sentiment analysis, aspect extraction, etc.), opinion summarization, or the task of
automatically creating a textual summary of opinions found in multiple reviews, aims
to help users access content and improves their decision making. This thesis focuses
on different methods to generate opinion summaries given multiple reviews about a
target entity (e.g., a product or service). The task is challenging due to the absence of
large-scale datasets for supervised training, which is paramount to the recent success
of neural-based systems. In this thesis, we propose several methods to synthesize these
datasets, thereby making supervised training for opinion summarization feasible.
Firstly, we introduce a two-step process that creates synthetic datasets for opinion
summarization. Given a corpus of reviews, we first sample a review and pretend it
is a (pseudo-)summary. Then, we procure a list of reviews to pair with the summary.
We obtain these reviews by generating noisy versions of the summary. We propose
a summarization model which learns to denoise the input reviews and generate the
summary, motivated by how humans write opinion summaries by removing divergent
opinions from reviews. Extensive evaluation shows that our model brings substantial
improvements over unsupervised abstractive and extractive baselines.
To further reflect the diversity of opinions in naturally-occurring reviews, we incorporate content planning during synthetic dataset creation. For each pseudo-summary
sampled from the corpus, we automatically induce its content plan in the form of aspect
and sentiment distributions. We then sample reviews from the corpus using Dirichlet
distributions parameterized by the content plan, and controlling the variance accordingly. Experimental results show that our approach outperforms competitive models in
generating opinion summaries that capture opinion consensus.
In opinion summarization, the notion of salience in reviews largely depends on user
interest, therefore a generic summary may not satisfy the needs of all users, limiting
their ability to make decisions. Therefore, we extend opinion summarization to generating aspect-controllable summaries. Using a synthetic training dataset enriched with
aspect controllers of different granularity, we fine-tune a pre-trained language model
which allows the creation of generic and aspect-specific summaries by modifying aspect controllers during inference. Experiments show that our model achieves state of
the art and is able to generate personalized summaries
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
- …
