1,721,102 research outputs found

    A integração dos papéis qualia para redes semânticas

    Get PDF
    Dissertação (mestrado) - Universidade Federal de Santa Catarina, Centro de Comunicação e Expressão. Programa de Pós-Graduação em Linguística.Este trabalho se desenvolve ao redor de dois eixos: o teórico, no qual se discutem alguns temas de ontologia filosófica e semântica lexical, como a individuação e a polissemia, o nominalismo e o realismo; e o técnico-prático, no qual se estabelece um conjunto de regras de inferência para ontologias computacionais, que chamaremos de QualiaNet, cujo objetivo é a integração das relações da estrutura qualia do léxico gerativo às árvores de inferência ou hierárquicas. São propostas regras como a seguinte: o constitutivo do formal de X é formal do contitutivo de X. Por fim, aplicamos esta mesma regra a uma parte das relações da Wordnet 2.0, por meio das fontes em PROLOG, e demonstramos que a QualiaNet não apenas incrementa coerência entre as relações semânticas como, também, pode ser utilizada para gerar novos arcos hiperonímicos, os quais são, como se sabe, a espinha dorsal de qualquer rede semântica

    Folktale similarity based on ontological abstraction

    No full text
    This paper presents a method to compute similarity of folktales based on conceptual overlap at various levels of abstraction as defined in Dutch WordNet. The method is applied on a corpus of Dutch folktales and evaluated using a comparison to traditional folktale similarity analysis based on the Aarne–Thompson–Uther (ATU) classification system. Document similarity computed by the presented method is in agreement with traditional analysis for a certain amount of folktale pairs, but differs for other pairs. However, it can be argued that the current approach computes an alternative, data-driven type of similarity. Using WordNet instead of a domain-specific ontology or classification system ensures applicability of the method outside of the folktale domain

    Netflips: A Content-Based Literature Recommendation System

    No full text
    Recommendation systems for media can be roughly broken down into two categories: collaborative systems, which rely on similarities between users, and content-based systems, which rely on similarities between items available to be recommended. Most systems today are collaborative, including most book recommendation platforms, such as Kindle Unlimited or Goodreads. However, the capabilities of natural language processing techniques to automatically extract large amounts of information from texts makes a content-based approach for book recommendation viable. This paper describes the implementation of a prototypical content-based book recommendation system, recommending a selection of books from Project Gutenberg’s public domain corpus. Based on preliminary results from a handful of informal test cases, this system was at least somewhat successful, providing satisfactory recommendations more than half of the time. It struggles with genre and subject matter alignment, though this could be mitigated through the incorporation of more abstract NLP techniques (such as topic modeling) into the feature extraction process. The fact that it is successful at all, while still being fairly rudimentary, indicates that content-based recommendation systems for literature have real potential

    Adversarial Learning for Bias Mitigation in Machine Translation

    No full text
    Natural language processing (NLP) systems often contain significant biases regarding sensitive attributes, such as gender, that worsen system performance and perpetuate harmful stereotypes. Recent research has found that adversarial neural networks can help to mitigate bias in word embeddings on the task of completing analogies, without impairing model performance. This model-agnostic method, which requires no cumbersome data modifications, aims to mitigate both the biases present in datasets and those amplified during training. However, this strategy still needs further development for use in downstream applications, particularly for use with large language models and language tasks in which gender or another protected variable like gender must be deduced from the data itself. To that end, this work proposes two methods, one based on the structure of sentence encodings and one based on the use of gendered pronouns, to measure gender representation in machine translation. It then presents an adversarial learning framework that uses these measures to mitigate gender bias in English-French translation on the language model T5. I found that this method successfully mitigated gender bias in the model’s translated output with minimal effect on translation quality. The results suggest that adversarial learning is a promising technique for use with large language models and complex, realistic applications

    Interrogating Computational Agent-Based Models of Diachronic Linguistic Change

    No full text
    In this work, we explore different approaches to the agent based model of diachronic linguistic change; in particular, we explore the simulation of consonant chain shifts, taking the well­ studied example of Grimm’s Law observed in Germanic languages as our case study. We present a particularly novel combination of continuous­ space representation and agent-­based modeling, showing that vector representations can be used to model constraints on phonological change. We successfully produce the sound shifts observed in Grimm’s Law under both deterministic and stochastic modeling constraints, and observe dynamics that warrant further investigation, observing s­-curve dynamics similar to those observed by field linguists. In particular, we show that under our modeling constraints, the emergence of the full chain of Grimm’s laws’ shifts is subject to variation by model parameters, presenting a novel paradigm under which to investigate the emergence and propagation of sound shifts

    E-bonics: An Analysis of AAVE's Context on an Electronic Social Media Platform

    No full text
    African American Vernacular English (AAVE) is a dialect utilized by a large percentage of the African American population. Today, AAVE's usage is stronger than ever, with its presence easily recognizable in social media comments and other online platforms. This phenomenon, defined as Mock AAVE, grows stronger and stronger in its employment among digital spaces. The observation of these occurrences led to a series of questions that this paper sought to answer: does the usage of AAVE qualities in text comments lead to more internet engagement and where in online spaces do these comments tend to cluster? The approach here details exploring a unique binary text classification system, one that takes into account AAVE vocabulary and a \Southern Similarity Index" inspired by the linguistic origins under the AAVE Dialect Hypothesis. This required the original contribution of Southern and non- Southern English datasets alongside a Southern American English classifier. After comparing a model with the contribution of the Southern Similarity Index to a model without that feature, it was shown that the performances of the two classifiers were comparable in the current methodology. This model was then run on a manually curated YouTube dataset that contains a subset of popular Black and non-Black content creators. The result of running this classifier yielded results that showed no inherent \comment reply" or \comment like" bias for comments classified as AAVE

    No Neutral Ground: The Language of "Us" and "Them" in Political Media Rhetoric

    No full text
    American political rhetoric has long been a battleground of "us" versus "them". In this thesis, I use computational methods of text analysis to examine news media during the 2008, 2012, and 2016 election periods, examining how language has been used in public conversation to communicate and reinforce division between political parties. Specifically, this thesis addresses the following questions: (1) Do media outlets use different vocabularies to describe the political party whose ideology aligns with their own, compared to the party on the opposite end of the ideological spectrum? (2) If so, do these different vocabularies also reflect a deeper division along ingroup and outgroup lines? and (3) Do these patterns change over time? I examine these questions using several distinct but complementary methods of textual analysis. First, I examine the extent to which lexical divides are present in the words and topics associated with different political parties. Next, I look at usages of emotion language and language abstractness, evaluating whether divisions are also reinforced semantically, in ways that communicate ingroup and outgroup boundaries. Ultimately, these different methods of analysis work together to paint a broad and multifaceted picture of political polarization in recent years

    Hyper-Correction and the Use of Pronoun Conjunctions in American English

    No full text
    This paper investigates how speakers of American English coordinate pronouns in sentences when one of those pronouns is a first person pronoun, focusing on hypercorrection. We found that the frequency in which those phrases are found in the language is important, but other factors were also present, such as experience with prescriptivism, language change, and deviation from prescriptive rules. For this paper, we used both the COCA corpus and an acceptability survey to better understand American English speakers’ attitudes towards different pronoun coordination types. Both corpus and survey data are necessary to get a full picture of how people use and judge language. This study helps us understand how grammar rules are changing in American English, while reflecting how meta-linguistic processes shape our use and acceptance of different utterances

    What Seems to be the Problem? Stigmatizing Language in Patient Medical Notes

    No full text
    Stigmatizing language in medical notes can prevent a patient from acquiring proper treatment. Reading medical notes containing biased language can influence subsequent clinicians’ perception of a patient, further compounding a patient’s inability to receive adequate care. Thus, there is a clear need to correct patient notes to eliminate stigmatizing language. Prior work involving stigmatizing language in medical notes has largely remained qualitative where clinicians and researchers manually analyzed notes for stigmatizing keywords. Our work utilized a computational approach to obtain a more robust set of stigmatizing keywords. We created contextual word embeddings from BERT-based and BioBERT-based models that are trained on free-text patient-oriented clinical data. These state-of-the-art models allowed us to develop word vector representations, from which we identified 30 new stigmatizing keywords. We then complete a thorough analysis to build a grammar structure that categorizes stigmatizing keywords according to the ways they induce stigma and better understand the syntactical environments in which these keywords occur. Following our analysis, we developed a model called MedStiLE (Medical note Stigmatizing Language Editor) that utilizes the grammar structure and constituency parsing to edit notes containing the stigmatizing keywords to be non-stigmatizing. We conducted an evaluation to test the efficacy of MedStiLE using human raters and found that it significantly reduced stigma in notes. This research provides various novel insights in terms of methodology and results that can help shape future works involving the intersection of language and healthcare
    corecore