1,720,963 research outputs found
SNAP judgments: A small N acceptability paradigm (SNAP) for linguistic acceptability judgments: Online Appendices
‘z-bad’ is the average z-score for the hypothesized ‘bad’ option. ‘z-good’ is the average z-score for the hypothesized good option. ‘Z.diff’ is the difference between z-good and z-bad and is the effect size. Beta is the estimate from the linear mixed-effects model, which has a standard error ‘SE’ and a t-value ‘t’. ‘χ²’ is the chi-squared value comparing the full model to an intercept-only model, and ‘χ² p’ is the p-value obtained by that comparison. Simple ‘p’ is just the p-value calculated using the t-value. Pred is TRUE if the effect goes in the significant direction. Sig is TRUE if there is a significant effect
SNAP judgments: A small N acceptability paradigm (SNAP) for linguistic acceptability judgments
While published linguistic judgments sometimes differ from the judgments found in large-scale formal experiments with naive participants, there is not a consensus as to how often these errors occur nor as to how often formal experiments should be used in syntax and semantics research. In this article, we first present the results of a large-scale replication of the Sprouse et al. 2013 study on 100 English contrasts randomly sampled from Linguistic Inquiry 2001–2010 and tested in both a forced-choice experiment and an acceptability rating experiment. Like Sprouse, Schütze, and Almeida, we find that the effect sizes of published linguistic acceptability judgments are not uniformly large or consistent but rather form a continuum from very large effects to small or nonexistent effects. We then use this data as a prior in a Bayesian framework to propose a small n acceptability paradigm for linguistic acceptability judgments (SNAP Judgments). This proposal makes it easier and cheaper to obtain meaningful quantitative data in syntax and semantics research. Specifically, for a contrast of linguistic interest for which a researcher is confident that sentence A is better than sentence B, we recommend that the researcher should obtain judgments from at least five unique participants, using at least five unique sentences of each type. If all participants in the sample agree that sentence A is better than sentence B, then the researcher can be confident that the result of a full forced-choice experiment would likely be 75% or more agreement in favor of sentence A (with a mean of 93%). We test this proposal by sampling from the existing data and find that it gives reliable performance.*American Society for Engineering Education. National Defense Science and Engineering Graduate Fellowshi
Large-scale evidence of dependency length minimization in 37 languages
Explaining the variation between human languages and the constraints on that variation is a core goal of linguistics. In the last 20 y, it has been claimed that many striking universals of cross-linguistic variation follow from a hypothetical principle that dependency length—the distance between syntactically related words in a sentence—is minimized. Various models of human sentence production and comprehension predict that long dependencies are difficult or inefficient to process; minimizing dependency length thus enables effective communication without incurring processing difficulty. However, despite widespread application of this idea in theoretical, empirical, and practical work, there is not yet large-scale evidence that dependency length is actually minimized in real utterances across many languages; previous work has focused either on a small number of languages or on limited kinds of data about each language. Here, using parsed corpora of 37 diverse languages, we show that overall dependency lengths for all languages are shorter than conservative random baselines. The results strongly suggest that dependency length minimization is a universal quantitative property of human languages and support explanations of linguistic variation in terms of general properties of human information processing.United States. Dept. of Defense. National Defense Science & Engineering Graduate Fellowship Progra
Wordform Similarity Increases With Semantic Similarity: An Analysis of 100 Languages
Although the mapping between form and meaning is often regarded as arbitrary, there are in fact well-known constraints on words which are the result of functional pressures associated with language use and its acquisition. In particular, languages have been shown to encode meaning distinctions in their sound properties, which may be important for language learning. Here, we investigate the relationship between semantic distance and phonological distance in the large-scale structure of the lexicon. We show evidence in 100 languages from a diverse array of language families that more semantically similar word pairs are also more phonologically similar. This suggests that there is an important statistical trend for lexicons to have semantically similar words be phonologically similar as well, possibly for functional reasons associated with language learning.Eunice Kennedy Shriver National Institute of Child Health and Human Development (U.S.) (Grant F32HD070544
Accommodating Presuppositions Is Inappropriate in Implausible Contexts
According to one view of linguistic information (Karttunen, 1974; Stalnaker, 1974), a speaker can convey contextually new information in one of two ways: (a) by asserting the content as new information; or (b) by presupposing the content as given information which would then have to be accommodated. This distinction predicts that it is conversationally more appropriate to assert implausible information rather than presuppose it (e.g., von Fintel, 2008; Heim, 1992; Stalnaker, 2002). A second view rejects the assumption that presuppositions are accommodated; instead, presuppositions are assimilated into asserted content and both are correspondingly open to challenge (e.g., Gazdar, 1979; van der Sandt, 1992). Under this view, we should not expect to find a difference in conversational appropriateness between asserting implausible information and presupposing it. To distinguish between these two views of linguistic information, we performed two self-paced reading experiments with an on-line stops-making-sense judgment. The results of the two experiments—using the presupposition triggers the and too—show that accommodation is inappropriate (makes less sense) relative to non-presuppositional controls when the presupposed information is implausible but not when it is plausible. These results provide support for the first view of linguistic information: the contrast in implausible contexts can only be explained if there is a presupposition-assertion distinction and accommodation is a mechanism dedicated to reasoning about presuppositions
Don’t Underestimate the Benefits of Being Misunderstood
Being a nonnative speaker of a language poses challenges. Individuals often feel embarrassed by the errors they make when talking in their second language. However, here we report an advantage of being a nonnative speaker: Native speakers give foreign-accented speakers the benefit of the doubt when interpreting their utterances; as a result, apparently implausible utterances are more likely to be interpreted in a plausible way when delivered in a foreign than in a native accent. Across three replicated experiments, we demonstrated that native English speakers are more likely to interpret implausible utterances, such as “the mother gave the candle the daughter,” as similar plausible utterances (“the mother gave the candle to the daughter”) when the speaker has a foreign accent. This result follows from the general model of language interpretation in a noisy channel, under the hypothesis that listeners assume a higher error rate in foreign-accented than in nonaccented speech.National Science Foundation (U.S.) (Award 1534318
A meta-analysis of syntactic priming in language production
We performed an exhaustive meta-analysis of 73 peer-reviewed journal articles on syntactic priming from the seminal Bock (1986) paper through 2013. Extracting the effect size for each experiment and condition, where the effect size is the log odds ratio of the frequency of the primed structure X to the frequency of the unprimed structure Y, we found a robust effect of syntactic priming with an average weighted odds ratio of 1.67 when there is no lexical overlap and 3.26 when there is. That is, a construction X which occurs 50% of the time in the absence of priming would occur 63% if primed without lexical repetition and 77% of the time if primed with lexical repetition. The syntactic priming effect is robust across several different construction types and languages, and we found strong effects of lexical overlap on the size of the priming effect as well as interactions between lexical repetition and temporal lag and between lexical repetition and whether the priming occurred within or across languages. We also analyzed the distribution of p-values across experiments in order to estimate the average statistical power of experiments in our sample and to assess publication bias. Analyzing a subset of experiments in which the primary result of interest is whether a particular structure showed a priming effect, we did not find evidence of major p-hacking and the studies appear to have acceptable statistical power: 82%. However, analyzing a subset of experiments that focus not just on whether syntactic priming exists but on how syntactic priming is moderated by other variables (such as repetition of words in prime and target, the location of the testing room, and the memory of the speaker), we found that such studies are, on average, underpowered with estimated average power of 53%. Using a subset of 45 papers from our sample for which we received raw data, we estimated subject and item variation and give recommendations for appropriate sample size for future syntactic priming studies. Keywords: Syntactic priming, Meta-analysis, Statistical powerAmerican Society for Engineering Education. National Defense Science and Engineering Graduate Fellowshi
Color naming across languages reflects color use
What determines how languages categorize colors? We analyzed results of the World Color Survey (WCS) of 110 languages to show that despite gross differences across languages, communication of chromatic chips is always better for warm colors (yellows/reds) than cool colors (blues/greens). We present an analysis of color statistics in a large databank of natural images curated by human observers for salient objects and show that objects tend to have warm rather than cool colors. These results suggest that the cross-linguistic similarity in color-naming efficiency reflects colors of universal usefulness and provide an account of a principle (color use) that governs how color categories come about. We show that potential methodological issues with the WCS do not corrupt information-theoretic analyses, by collecting original data using two extreme versions of the color-naming task, in three groups: the Tsimane’, a remote Amazonian hunter-gatherer isolate; Bolivian-Spanish speakers; and English speakers. These data also enabled us to test another prediction of the color-usefulness hypothesis: that differences in color categorization between languages are caused by differences in overall usefulness of color to a culture. In support, we found that color naming among Tsimane’ had relatively low communicative efficiency, and the Tsimane’ were less likely to use color terms when describing familiar objects. Color-naming among Tsimane’ was boosted when naming artificially colored objects compared with natural objects, suggesting that industrialization promotes color usefulness.National Science Foundation (U.S.) (Award 1534318
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
- …
