70 research outputs found

    Learning the easy way: the role of form similarity in language learning and processing

    Get PDF
    A key question about bilingual lexical access is whether lexical representations are activated selectively (within one language) or non-selectively (across languages). A great deal of the research addressing this question has focused on cognates, translation equivalents with the same or similar forms across languages. These studies show that, in non-native language (L2) processing, cognates are usually recognised and produced faster than non-cognates, suggesting that the form overlap between translation equivalents contributes to a processing advantage. Traditionally, the advantage is assumed to demonstrate on-line non-selective activation: For cognates, lexical activation stems from two sources rather than one, as is the case with non-cognates. This converging activation leads to their facilitated processing. However, some researchers argue that cognate facilitation could be due to differences in how cognates and non-cognates are learned, which leads to qualitatively different representations for cognates. In this thesis, I focus on learning-based explanations of cognate effects and ask to what extent these can account for cognate facilitation. First, I review findings of cognate effects in both comprehension and production, evaluating whether they are more consistent with learning-based or on-line accounts. Second, I explore whether neural language models trained on two languages exhibit cognate facilitation and use this to test learning-based hypotheses of the effect. Following this, I present two behavioural experiments investigating cognate effects in human non-native speakers. The first investigates how the language of instruction affects cognate facilitation in trilinguals to examine whether cognate effects only occur between languages involved in learning. The second tests whether bilinguals exhibit cognate effects in L2 prediction to shed light on whether predictive processing is language selective. Taken together, this thesis provides evidence that cognate facilitation can, in principle, be explained by learning. I argue that while non-selective activation may occur during on-line processing, it is not required to explain cognate effects. I call for a more nuanced approach to theories of bilingual lexical access that allow for a more flexible role for language selectivity

    A Model Comparison between Neural Architectures of Human Bilingual Sentence Processing

    Get PDF
    This work investigates phenomena related to human bilingual sentence processing in neural language models. We ask ourselves the question if and how the emergence of these phenomena depends on the model architecture. For this purpose, we train SRNs, LSTMs, and Transformers with different hidden layer sizes as bilingual- and monolingual language models. We test these models on three phenomena that have been shown to emerge in at least one of the architectures in the literature. We refer to them as reading time prediction; an agreement between monolingual vs. bilingual reading with models trained on monolingual vs. bilingual data, the cognate facilitation effect; a faster processing of form and meaning-similar words, and the grammaticality illusion; a preference for the ungrammatical version of a certain class of sentences that is reversed in some languages. Surprisingly, we found reading time prediction to depend not only on architecture and layer size, but also on the specific random initialization. Failure of reproduction of the effect was confirmed by the author, suggesting the original study to be due to coincidence. As for the cognate facilitation effect, we found it to be present in the SRN and LSTM, providing further evidence for its emergence in humans to be due to the cumulative frequency of cognates. The effect was found to decrease in magnitude for large layer sizes in the LSTM, which can be linked to the LSTM relying less on corpus frequency. However surprisingly, the effect was found to increase for small layer sizes in the SRN. We do not have an adequate explanation for this trend. Furthermore, it was not found in the Transformer, suggesting that Transformers exhibit less cross-linguistic transfer than the other architectures. The grammaticality illusion was found to be present in the SRN, but not in the LSTM and Transformer. This provides further evidence for the effect to arise as a result of short-distance language statistics rather than universal working memory constraints. The effect was found to stay fairly constant over layer size, syntactic linguistic transfer to be small. Furthermore, the Transformer displayed a consistent preference for grammatical sentences, suggesting super-human syntactic proficiency on this particular task

    Modeling color naming in bilinguals: Computational mechanisms of crosslinguistic influence

    No full text
    When bilingual speakers name stimuli such as colors or objects, their naming patterns can differ from those of monolingual speakers. Three accounts have been proposed to explain these differences – conceptual change, online lexical coactivation, and L1 footprint – yet these have not been empirically evaluated against each other. In this study, we propose a novel computational cognitive model which operationalizes each of these proposals as a mechanism of crosslinguistic influence, such that we can study their individual and combined effects on the model’s behavior. We focus on the domain of color in which we model existing experimental data collected from Navajo and English monolinguals and Navajo–English bilinguals. Our color learning model extends a statistical learning procedure for mixture models to the acquisition of labelled categories, and achieves bilingual learning by maintaining two sets of color categories and associated color words, which are connected in varying ways according to the three crosslinguistic mechanisms. We test the combinations of mechanisms in a color naming task, and analyze the match between the naming patterns of the model and the differences between bilingual and monolingual human speakers. Our results suggest that gradual conceptual change following crosslinguistic transfer at the initial learning stage can best capture the observed differences in human color naming patterns. While lexical coactivation combined with initial transfer can account for some of the empirical data, this mechanism consistently performs less well than that of conceptual change

    Trees neural those:RNNs can learn the hierarchical structure of noun phrases

    Get PDF
    Humans use both linear and hierarchical representations in language processing, and the exact role of each has been debated. One domain where hierarchical processing is important is noun phrases. English noun phrases have a fixed order of prenominal modifiers: demonstratives - numerals - adjectives (these two green vases). However, when English speakers learn an artificial language with postnominal modifiers, instead of reproducing this linear order they preserve the distance between each modifier and the noun (vases green two these). This has been explained by a hierarchical homomorphism bias. Here, we investigate whether RNNs exhibit this bias. We pre-train one linear and two hierarchical models on English and expose them to a small artificial language. We then test them on noun phrases from a study with humans and find that only the hierarchical models can exhibit the bias, supporting the idea that homomorphic word order preferences arise from hierarchical, and not linear relations

    Multilingual Acoustic Word Embedding Models for Processing Zero-Resource Languages

    Get PDF
    Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in “zero-resource” speech search, indexing and discovery systems. Here we propose to train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zeroresource languages. For this transfer learning approach, we consider two multilingual recurrent neural network models: a discriminative classifier trained on the joint vocabularies of all training languages, and a correspondence autoencoder trained to reconstruct word pairs. We test these using a word discrimination task on six target zero-resource languages. When trained on seven well-resourced languages, both models perform similarly and outperform unsupervised models trained on the zero-resource languages. With just a single training language, the second model works better, but performance depends more on the particular training–testing language pair
    corecore