70 research outputs found
Learning the easy way: the role of form similarity in language learning and processing
A key question about bilingual lexical access is whether lexical representations are
activated selectively (within one language) or non-selectively (across languages). A
great deal of the research addressing this question has focused on cognates, translation
equivalents with the same or similar forms across languages. These studies show that,
in non-native language (L2) processing, cognates are usually recognised and produced
faster than non-cognates, suggesting that the form overlap between translation equivalents
contributes to a processing advantage. Traditionally, the advantage is assumed
to demonstrate on-line non-selective activation: For cognates, lexical activation stems
from two sources rather than one, as is the case with non-cognates. This converging
activation leads to their facilitated processing. However, some researchers argue that
cognate facilitation could be due to differences in how cognates and non-cognates are
learned, which leads to qualitatively different representations for cognates.
In this thesis, I focus on learning-based explanations of cognate effects and ask
to what extent these can account for cognate facilitation. First, I review findings of
cognate effects in both comprehension and production, evaluating whether they are more
consistent with learning-based or on-line accounts. Second, I explore whether neural
language models trained on two languages exhibit cognate facilitation and use this to
test learning-based hypotheses of the effect. Following this, I present two behavioural
experiments investigating cognate effects in human non-native speakers. The first
investigates how the language of instruction affects cognate facilitation in trilinguals to
examine whether cognate effects only occur between languages involved in learning.
The second tests whether bilinguals exhibit cognate effects in L2 prediction to shed
light on whether predictive processing is language selective. Taken together, this thesis
provides evidence that cognate facilitation can, in principle, be explained by learning. I
argue that while non-selective activation may occur during on-line processing, it is not
required to explain cognate effects. I call for a more nuanced approach to theories of
bilingual lexical access that allow for a more flexible role for language selectivity
A Model Comparison between Neural Architectures of Human Bilingual Sentence Processing
This work investigates phenomena related to human bilingual sentence processing in neural language models. We ask ourselves the question if and how the emergence of these phenomena depends on the model architecture. For this purpose, we train SRNs, LSTMs, and Transformers with different hidden layer sizes as bilingual- and monolingual language models. We test these models on three phenomena that have been shown to emerge in at least one of the architectures in the literature. We refer to them as reading time prediction; an agreement between monolingual vs. bilingual reading with models trained on monolingual vs. bilingual data, the cognate facilitation effect; a faster processing of form and meaning-similar words, and the grammaticality illusion; a preference for the ungrammatical version of a certain class of sentences that is reversed in some languages.
Surprisingly, we found reading time prediction to depend not only on architecture and layer size, but also on the specific random initialization. Failure of reproduction of the effect was confirmed by the author, suggesting the original study to be due to coincidence. As for the cognate facilitation effect, we found it to be present in the SRN and LSTM, providing further evidence for its emergence in humans to be due to the cumulative frequency of cognates. The effect was found to decrease in magnitude for large layer sizes in the LSTM, which can be linked to the LSTM relying less on corpus frequency. However surprisingly, the effect was found to increase for small layer sizes in the SRN. We do not have an adequate explanation for this trend. Furthermore, it was not found in the Transformer, suggesting that Transformers exhibit less cross-linguistic transfer than the other architectures. The grammaticality illusion was found to be present in the SRN, but not in the LSTM and Transformer. This provides further evidence for the effect to arise as a result of short-distance language statistics rather than universal working memory constraints. The effect was found to stay fairly constant over layer size, syntactic linguistic transfer to be small. Furthermore, the Transformer displayed a consistent preference for grammatical sentences, suggesting super-human syntactic proficiency on this particular task
Recommended from our members
Bilingual Sentence Processing: when Models Meet Experiments
Although sentence comprehension and production are increasingly often studied by combining computational modeling and human experiments, this approach remains mostly restricted to studies of monolingual or first-language (L1) processing. There are currently only very few sentence-level computational models of second-language (L2) or bilingual processing (Frank, 2021). This lack of computational specifications can hamper further progress in bilingualism research. Moreover, better understanding of bilingual processing will give more insights into more general mechanisms such as cognitive control processes involved while switching languages (Luk et al., 2012). Our symposium aims to bring together researchers from different labs and with different research traditions, working on the intersection of models and experiments in bilingual sentence processing.
The symposium has four talks, by Edith Kaan (associate professor, specializing in psycholinguistics of bilingualism), Yung Han Khoe (PhD student, working on models of bilingual sentence production), Lin Chen (research associate with an expertise in reading processes), and a joint talk by Irene Winther (PhD student working on bilingual sentence processing) and Yevgen Matusevych (research associate in computational cognitive science of language). Finally, we will have a panel discussion to suggest how models could be challenged by experimental data, and provide new explanatory mechanisms. This discussion will be moderated by Xavier Hinaut (research scientist in computational neuroscience) and Stefan Frank (associate professor in computational psycholinguistics)
Learning constructions from bilingual exposure: Computational studies of argument structure acquisition
Modeling color naming in bilinguals: Computational mechanisms of crosslinguistic influence
When bilingual speakers name stimuli such as colors or objects, their naming patterns can differ from those of monolingual speakers. Three accounts have been proposed to explain these differences – conceptual change, online lexical coactivation, and L1 footprint – yet these have not been empirically evaluated against each other. In this study, we propose a novel computational cognitive model which operationalizes each of these proposals as a mechanism of crosslinguistic influence, such that we can study their individual and combined effects on the model’s behavior. We focus on the domain of color in which we model existing experimental data collected from Navajo and English monolinguals and Navajo–English bilinguals. Our color learning model extends a statistical learning procedure for mixture models to the acquisition of labelled categories, and achieves bilingual learning by maintaining two sets of color categories and associated color words, which are connected in varying ways according to the three crosslinguistic mechanisms. We test the combinations of mechanisms in a color naming task, and analyze the match between the naming patterns of the model and the differences between bilingual and monolingual human speakers. Our results suggest that gradual conceptual change following crosslinguistic transfer at the initial learning stage can best capture the observed differences in human color naming patterns. While lexical coactivation combined with initial transfer can account for some of the empirical data, this mechanism consistently performs less well than that of conceptual change
Recommended from our members
Analyzing and modeling free word associations
Human free association (FA) norms are believed to reflect the strength of links between words in the lexicon of an average speaker. Large-scale FA norms are commonly used as a data source both in psycholinguistics and in computational modeling. However, few studies aim to analyze FA norms themselves, and it is not known what are the most important factors that guide speakers’ lexical choices in the FA task. Here, we first provide a statistical analysis of a large-scale data set of English FA norms. Second, we argue that such analysis can inform existing computational models of semantic memory, and present a case study with the topic model to support this claim. Based on our analysis, we provide the topic model with dictionary-based knowledge about word synonymy/antonymy, and demonstrate that the resulting model predicts human FA responses better than the topic model without this information
Recommended from our members
Modeling Sentence Processing Effects in Bilingual Speakers: A Comparison of Neural Architectures
Neural language models are commonly used to study language processing in human speakers, and several studies trained such models on two languages to simulate bilingual speakers. Surprisingly, no work systematically evaluates different neural architectures on bilingual speakers’ data, despite the abundance of such studies in the monolingual domain. In this work, we take the first step in this direction. We train three neural architectures (SRN, LSTM, and Transformer) on Dutch and English data and evaluate them on two data sets from experimental studies. Our goal is to investigate which architectures can reproduce the cognate facilitation effect and grammaticality illusion observed in bilingual speakers. While all three architectures can correctly predict the cognate effect, only the SRN succeeds at the grammaticality illusion. We additionally show how the observed patterns change as a function of the models’ hidden layer size, a hyperparameter that we argue may be more important in bilingual models
Trees neural those:RNNs can learn the hierarchical structure of noun phrases
Humans use both linear and hierarchical representations in language processing, and the exact role of each has been debated. One domain where hierarchical processing is important is noun phrases. English noun phrases have a fixed order of prenominal modifiers: demonstratives - numerals - adjectives (these two green vases). However, when English speakers learn an artificial language with postnominal modifiers, instead of reproducing this linear order they preserve the distance between each modifier and the noun (vases green two these). This has been explained by a hierarchical homomorphism bias. Here, we investigate whether RNNs exhibit this bias. We pre-train one linear and two hierarchical models on English and expose them to a small artificial language. We then test them on noun phrases from a study with humans and find that only the hierarchical models can exhibit the bias, supporting the idea that homomorphic word order preferences arise from hierarchical, and not linear relations
Multilingual Acoustic Word Embedding Models for Processing Zero-Resource Languages
Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in “zero-resource” speech search, indexing and discovery systems. Here we propose to train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zeroresource languages. For this transfer learning approach, we consider two multilingual recurrent neural network models: a discriminative classifier trained on the joint vocabularies of all training languages, and a correspondence autoencoder trained to reconstruct word pairs. We test these using a word discrimination task on six target zero-resource languages. When trained on seven well-resourced languages, both models perform similarly and outperform unsupervised models trained on the zero-resource languages. With just a single training language, the second model works better, but performance depends more on the particular training–testing language pair
- …
