1,721,103 research outputs found

    Sources and Targets in Kuteva et al. 2019

    No full text
    This dataset is based on examples found in Kuteva et al. 2019: Kuteva, Tania, Bernd Heine, Bo Hong, Haiping Long, Heiko Narrog, and Seongha Rhee. 2019. World Lexicon of Gramaticalization (2nd ed.). Cambridge: Cambridge University Press. Kuteva et al.’s World Lexicon of Grammaticalization (2019) is an inventory of examples of morphological reanalysis observed across a sample of over 900 languages. The goal of Kuteva et al.’s inventory is to represent grammaticalization changes that are documented in multiple languages. The examples are cataloged as types defined by Source to Target shifts, such as Ablative > Partitive, that are attested in two or more languages of the sample. The inventory lists 526 Source > Target types. The purpose of this dataset is to explore how many Sources also serve as Targets, and to provide a broad semantic classification of the Sources. The classification is intended only for general description of patterns, and does not represent a precise assignment to mutually exclusive classes. Its purpose is to give a qualitative overview and thus does not lend itself to further quantitative analysis. This dataset is the basis for the analysis in Section 4 of this publication: Janda, Laura A. To Appear. “Morphological reanalysis: recycling old form to new function”, as part of Volume 3 Morphology & Syntax, Part 1 Morphology of The Wiley Blackwell Companion to Diachronic Linguistics, edited by Edith Aldridge, Anne Breitbarth, Katalin É. Kiss, Adam Ledgeway, Joe Salmons, and Alexandra Simonenko.</p

    Replication Data for: Contextually determined or semantically distinct? The competition between instrumental, long form nominative and short form nominative in Russian predicate adjectives

    No full text
    Dataset description This post provides the data and R scripts for analysis of data on the variation between long form nominative, short form nominative, and instrumental case in Russian predicate adjectives in sentences containing an overt copula verb. We analyze the various factors associated with the choice of form of the adjective.This is the abstract of the article: Based on data from the syntactic subcorpus of the Russian National Corpus, we undertake a quantitative analysis of the competition between Russian predicate adjectives in the instrumental (e.g., pustym ‘empty’), the long form nominative (e.g., pustoj ‘empty’), and the short form nominative (e.g., pust ‘empty’). It is argued that the choice of adjective form is partly determined by the context. Four (nearly) categorical rules are proposed based on the following contextual factors: the form of the copula verb, the presence/absence of a complement, and the nature of the subject of the sentence. At the same time, a “space of competition” is identified, where all three adjective forms are attested. It is hypothesized that within the space of competition, the three forms are recruited to convey different meanings, and it is argued that our analysis lends support to the traditional idea that the short form nominative is closely related to verbs. Our findings are furthermore compatible with the idea that the short form nominative expresses temporary states, rather than inherent permanent characteristics.</p

    Replication Data for: Typology of reduplication in Russian: constructions within and beyond a single clause

    No full text
    We analyze repetition in Russian from the perspective of the Russian Constructicon which represents over 2200 grammatical constructions described in terms of anchors (fixed elements) and slots (for various filler elements) and fully annotated for their syntactic and semantic characteristics. The Russian Constructicon facilitates the first large-scale investigation of reduplication across a representative sample of an entire language, enabling us to map out a typology invoking these and other factors in the context of Construction Grammar. Our data on repetitions includes 118 constructions tagged the Russian Constructicon for Reduplication, meaning that repetition occurs within a clause, and 28 entries tagged as Discourse “Echo” Constructions because they require the repetition of a word or phrase from a previous clause (often provided by an interlocutor). Five constructions carry both tags. We propose a theoretical expansion of the definition of reduplication to include the Discourse “Echo” type, arguing that constructions are not limited to a single clause or even to a single speaker. Our typology further explores the distribution of various formal and semantic factors observed in constructions with repetition and compares them with both previous typological research on reduplication and their distribution across the entire Russian Constructicon. Despite the fact that Russian does not use reduplication as a productive grammatical marker, we argue that reduplication is widespread and systematic in Russian

    Replication Data for: The long and the short of it: Russian predicate adjectives with zero copula

    No full text
    Description of Dataset This is a study of examples of Russian predicate adjectives in clauses with zero-copula present tense, where the adjective is a short form (SF) or a long form nominative (LF). The data was collected in 2022 from SynTagRus (https://universaldependencies.org/treebanks/ru_syntagrus/index.html), the syntactic subcorpus of the Russian National Corpus (https://ruscorpora.ru/new/). The data merges the results of several searches conducted to extract examples of sentences with long form and short form adjectives in predicate position, as identified by the corpus. The examples were imported to a spreadsheet and annotated manually, based on the syntactic analyses given in the corpus. For present tense sentences with no copula (Река спокойна or Река спокойная), it was necessary to search for an adjective as the top (root) node in the syntactic structure. The syntactic and morphological categories used in the corpus are explained here: https://ruscorpora.ru/page/instruction-syntax/. In order for the R code to run from these files, one needs to set up an R project with the data files in a folder named "data" and the R markdown files in a folder named "scripts". Method: Logistic regression analysis of corpus data carried out in R (R version 4.2.3 (2023-03-15)-- "Shortstop Beagle" Copyright (C) 2023 The R Foundation for Statistical Computing) and documented in an .Rmd file.Publication Abstract The present article presents an empirical investigation of the choice between so-called long (e.g., prostoj ‘simple’) and short forms (e.g., prost ‘simple’) of predicate adjectives in Russian based on data from the syntactic subcorpus of the Russian National Corpus. The data under scrutiny suggest that short forms represent the dominant option for predicate adjectives. It is proposed that long forms are descriptions of thematic participants in sentences with no complement, while short forms may take complements and describe both participants (thematic and rhematic) and situations. Within the “space of competition” where both long and short forms are well attested, it is argued that the choice of form to some extent depends on subject type, gender/number, and frequency. On the methodological level, the approach adopted in the present study may be extended to other cases of competition in morphosyntax. It is suggested that one should first “peel off” contexts where (nearly) categorical rules are at work, before one undertakes a statistical analysis of the “space of competition”.</p

    Replication Data for: Going Beyond Words: Engaging Grammar for Insights into Political Discourse

    No full text
    Dataset description: This dataset contains data in connection with a selection of three of Putin's speeches from 2023 and 2024. The related book chapter also includes analysis of data from Putin's speeches in 2022, and that data is available here: Obukhova, A. (2022). Replication Data for: the Case for Case in Putin’s Speeches. https://doi.org/10.18710/APDMDZ. DataverseNO, V2.Abstract for related publication: We present Keymorph Analysis as method to reveal the role of grammar in political discourse, demonstrating how this method can be used to gain a more in-depth understanding of language in discourse and securitization analyses. We present a longitudinal study of Putin’s speeches, comparing his use of grammatical case to usual grammatical behavior in Russian, and tracking changes from the beginning of the full-scale invasion of Ukraine in early 2022 through the end of 2024.</p

    Replication Data for: Understanding ‘many’ through the lens of Ukrainian багато

    No full text
    Dataset description: The General Regionally Annotated Corpus of Ukrainian (GRAC, Shvedova et al. 2017-2024, uacorpus.org) was consulted to collect data for further analysis concerning the distribution of Singular vs. Plural verb forms in the target bahato construction. GRAC is a Sketch Engine corpus of over 1.8 billion words, representing texts from over 30,000 authors created between 1816 and 2023. This corpus is designed to serve as source material for linguistic research on Standard Ukrainian. Our data was collected during the month of February 2024. We extracted and annotated 28,491 examples of the bahato construction. An additional set of examples was collected from the Russian National Corpus (ruscorpora.ru) during the month of August 2024 to provide comparison with the Russian mnogo construction. For this purpose, 6,612 examples were extracted and annotated for word order and Singular vs. Plural verb agreement. Both the Ukrainian and the Russian data are included in this dataset, along with the R scripts used to analyze this data. Article abstract: We reveal an ongoing language change in Ukrainian involving a construction with a subject comprised of the indefinite quantifier багато ‘many’ modifying a noun phrase in the Genitive Plural. Number agreement on the verb varies, allowing both Singular (in 69.1% of attestations) and Plural (in 30.9% of attestations). Based on statistical analysis of corpus data, we investigate the influence of the factors of year of creation, word order of subject and verb, and animacy of the subject on the choice of verb number. We find that, while all combinations of word order and animacy are robustly attested, VS word order and inanimate subjects tend to prefer Singular, whereas SV word order and animate subjects tend to prefer Plural. Since about the 1950s, the proportion of Plural has been increasing, overtaking Singular in the current decade. We propose that this Singular vs. Plural variation is motivated by the human embodied experience of construing a group of items as either a homogeneous mass (and therefore Singular) or a multiplicity of individuals (and therefore Plural). This proposal is supported by the identification of micro-constructions that prefer Singular and show reduced individuation of human beings

    Replication Data for: Davvisámi earutkeahtes oamasteapmi

    No full text
    This data shows the correlation analysis for our study with this description: On the basis of corpus data (9.5M words 1997-2010) we claim that North Saami is developing a grammatical distinction between alienable and inalienable possession. In previous work we documented a language change in North Saami in which the possessive suffix (“SOG”) as in girjji-id-easkka [book-ACC.PL-3PL] ‘their books’ is being replaced by an analytic construction with the reflexive genitive ieža-form, as in iežaska girjjiid ‘their books’. According to typologists, alienable/inalienable distinctions arise primarily in small languages where a language change takes place, and inalienability is marked by the synthetic construction. North Saami possessive constructions comport with these features, and SOG tends to mark inalienable possession, as opposed to the more neutral and widespread ieža-form. Statistical analysis shows that word frequency cannot account for the distribution of SOG vs. ieža-form, justifying focus on semantics. North Saami shows high frequency of SOG for kinship and body part nouns associated with inalienability cross-linguistically, but in addition extends this category to words for friends. A new finding is the strong presence of SOG with words for products and experiences, and additionally words connected with identity and way of life. SOG is productive lexically and morphologically, and used in multiple collocations

    Replication Data for: A network of allostructions: quantified subject constructions in Russian

    No full text
    Data and R code are provided for statistical analysis of approximately 39,000 corpus examples of predicate agreement in constructions with quantified subjects in Russian. The analysis indicates that these constructions constitute a network of constructions (“allostructions”) with various preferences for singular or plural agreement. Factors pull in different directions, and we observe a relatively stable situation in the face of variation. We present an analysis of a multidimensional network of allostructions in Russian, thus contributing to our understanding of allostructional relationships in Construction Grammar. With regard to historical linguistics, language stability is an understudied field. We illustrate an interplay of divergent factors that apparently resists language change. The syntax of numerals and other quantifiers represents a notoriously complex phenomenon of the Russian language. Our study sheds new light on the contributions of factors that favor singular or plural agreement in sentences with quantified subjects

    Replication Data for: Seeing from without, seeing from within: aspectual differences between Spanish and Russian

    No full text
    This is the data that serves as the basis for an article comparing the grammatical category of aspect in Spanish and Russian. Here is the abstract of the article: Linguistic categories such as aspect are not identical across languages, and cross-linguistic differences can reveal differences in construal and conceptual categorization, which are key concepts in cognitive linguistics. Spanish-Russian parallel data diverge in situations where Spanish uses a Perfective Past tense form, while the Russian translation equivalent is an Imperfective Past tense form. We classify examples of aspectual mismatch according to grammatical constructions and language-specific facts. We find this mismatch in contexts with overt expression of time periods, as well as situations in which a final temporal boundary either is expressed or can be inferred. We interpret this in terms of a difference in conceptualization: Spanish has a tendency to view time periods from without, interpreting them as bounded and thus Perfective, whereas Russian has a tendency to view time periods from within, interpreting them on the basis of their duration without reference to their boundaries and thus Imperfective

    Replication Data for: Looking into the Russian future

    No full text
    This dataset concerns the data for the article that covers the topic of future tense meanings in Russian. Abstract: The relationship between future time and future tense forms in Russian is complex. The forms traditionally attributed to the future tense in certain cases do not refer to future time. Those cases have been previously presented as a list and/or attributed to the sphere of modality. In this article, we suggest a data-driven approach applied to the spectrum of meanings of Russian future tense forms. We analyzed corpus data and discovered that 44% of perfective future forms and 22% of imperfective future forms do not unambiguously express future time meaning. Among the non-future time meanings that Russian future tense forms can express are Gnomic, Performative, Implicative, Hypothetical, Alternation, and Stable scenario. Furthermore, we propose that the meanings of the future tense constitute a radial category. Future time reference is the prototypical meaning of the future tense. The remaining meanings comprise extensions connected to the prototypical meaning. We describe the radial category with reference to Langacker’s (2008) model of tense and potentiality. Additionally, we explore the interaction of future tense and modality.</p
    corecore