1,720,967 research outputs found
INEL Enets Corpus
Corpus Citation
Shluinsky, Andrey; Khanina, Olesya; Wagner-Nagy, Beáta. 2024. INEL Enets Corpus. Version 1.0. Publication date 2024-11-30. https://hdl.handle.net/11022/0000-0007-FE1D-C. Archived at Universität Hamburg. In: The INEL corpora of indigenous Northern Eurasian languages. https://hdl.handle.net/11022/0000-0007-F45A-1
Corpus Description
The INEL Enets corpus has been created within the long-term INEL project ("Grammatical Descriptions, Corpora and Language Technology for Indigenous Northern Eurasian Languages"), 2016–2033.
The corpus includes texts recorded between 1962–2017 in both Enets lects – Forest Enets and Tundra Enets. The sources of the corpus (see more details in the user documentation, section 2.2) are:
Audio recordings done by Olesya Khanina, Maria Ovsjannikova, Andrey Shluinsky, Natalia Stoynova and Sergey Trubetskoy,
Legacy audio recordings done by Vera Bettu, Nina N. Bolina, Dar`ya S. Bolina, Zoya N. Bolina, Oksana E. Dobzhanskaya, Valentin Gusev, Eugene Helimski†, Kazimir I. Labanauskas†, Larisa Leisiö, Marina Lyublinskaya, Kaur Mägi, Viktor N. Pal`chin, Marina N. Pal`china, Irina P. Sorokina†, Anna Urmanchieva, Beáta Wagner-Nagy and possibly other people,
Published audio recordings,
Texts published by Dar`ya S. Bolina, Yaroslav A. Gluxij† and Vasilij A. Susekov†, Eugene Helimski†, Kazimir I. Labanauskas†, Tibor Mikola†, János Pusztay, Irina P. Sorokina†, Anna Urmanchieva,
Legacy manuscript transcriptions and self-transcriptions done and/or edited by Dar`ya S. Bolina, Galina S. Bolina, Zoya N. Bolina, Valentin Gusev, Eugene Helimski†, Kazimir I. Labanauskas†, Larisa Leisiö, Marina Lyublinskaya, Vasilij F. Ly`rmin†, Anton N. Pal`chin, Viktor N. Pal`chin, Ivan I. Silkin†, Irina P. Sorokina†, Natal`ya M. Tereščenko†, Anna Urmanchieva and possibly other people.
All texts in the corpus are provided with interlinear morpheme-by-morpheme glosses and translation into English and Russian. All texts for which the audio recordings were accessible are time-aligned with them. Video recordings are also included into the corpus if available.
Corpus size
Forest Enets: 541 texts, 41,396 sentences, 173,379 tokens
Tundra Enets: 137 texts, 12,737 sentences, 45,331 tokens
Total: 678 texts, 54,133 sentences, 218,710 tokens
Total duration of audio: 43 hours 26 minutes
Funding
The corpus has been produced in the context of the joint research funding of the German Federal Government and Federal States in the Academies’ Programme, with funding from the Federal Ministry of Education and Research and the Free and Hanseatic City of Hamburg. The Academies’ Programme is coordinated by the Union of the German Academies of Sciences and Humanities.
Preliminary glossing work included into this corpus was supported by Endangered Languages Documentation Programme (ELDP) and by Max Planck Institute for Evolutionary Anthropology (MPI-EVA). See more details on financial support in the documentation file below, section 1.6.
Contributions/Acknowledgements
Dozens of people and many institutions contributed to the corpus (see more details in the documentation file below, section 1.6). We are especially grateful to:
Enets speakers who generously shared their knowledge, especially those who spent many days working with us: Aleksandr S. Bolin†, Leonid D. Bolin†, Viktor N. Bolin, Nadezhda K. Bolina, Nina N. Bolina, Ekaterina S. Glibchenko, Gennadij A. Ivanov†, Irina P. Koshkaryova†, Valentina P. Nader, Lyudmila P. Novosyolova, Svetlana A. Roslyakova†, Ivan I. Silkin†, Nikolaj I. Silkin, Alevtina S. Silkina, Zoya A. Turutina, Tat`yana Ch. Yar,
In particular, Zoya N. Bolina and Viktor N. Pal`chin who also collaborated in ELDP project and extensively transcribed Enets recordings,
Natalia Stoynova, Sergey Trubetskoy and foremostly Maria Ovsjannikova who did recordings and transcriptions of Enets texts,
Institutions and private individuals who shared legacy data: the Institute for Linguistic Studies RAS, the Taymyr House of National Arts, the Dudinka branch of GTRK “Norilsk”; Dar`ya S. Bolina, Oksana E. Dobzhanskaya, Valentin Gusev, Larisa Leisiö, Viktor N. Pal`chin, Irina P. Sorokina†, Anna Urmanchieva,
Marina Lyublinskaya and Anna Urmanchieva who kindly permitted to include texts processed by them into the corpus,
Dar`ya S. Bolina who consulted a lot in the process of compilation of the corpus.
Searching the corpus
The corpus can be downloaded from the ZFDM Repository using the links provided below and browsed or searched locally using the EXMARaLDA software or, alternatively, ELAN.
Online search with Tsakorpus platform is available at https://inel.corpora.uni-hamburg.de/EnetsCorpus/search.
Remote search with EXMARaLDA is also possible without downloading all the files (see https://inel.corpora.uni-hamburg.de/portal/help/en/index.php#search).
See the user documentation (section 3) for details on transcription, annotation tiers and annotation tags.
Find further information and links on the Enets Corpus page at the INEL Resources portal: https://inel.corpora.uni-hamburg.de/portal/corpora/enets/
Morphosyntax of associated motion constructions in Akebu
The paper deals with associated motion constructions in Akebu (< Ghana-Togo Mountain < Kwa). Associated motion is expressed either by serial verb constructions with the verbs ‘go’ and ‘come’ or by andative and ventive morphological markers, which are etymologically clearly based on the same verbs. Morphological markers and serial verb constructions have a distribution that is close to complimentary, depending on verbal tense-aspect-modality-polarity grams and partially on morphological peculiarities of verbal forms. Despite the morphological difference between these two ways of expressing associated motion, they appear to be uniform syntactically
Verbal prefixes mà- and rà- in Susu and lexical features of verbal stems
L’article offre un analyse systématique des données disponibles concernant les deux préfixes verbaux les plus productifs de la langue susu, mà- et rà-. D’une part, chacun de ces préfixes a un ensemble de sens mutuellement liés, et le choix du sens dépend largement des caractéristiques sémantiques de la base verbale. D’autre part, chacun des préfixes manifeste de nombreux emplois lexicalisés.This paper presents a systematic analysis of the available data related to the two most productive verbal prefixes in Susu, mà- and rà-. On the one hand, each of the two prefixes has a number of semantically interrelated meanings, and the choice of a particular meaning depends, to a significant extent, on the lexical semantic features of the verbal stem. On the other hand, there are many lexicalized uses of both prefixes.В статье систематизированы данные об употреблении двух наиболее продуктивных глагольных приставок в языке сусу, mà- и rà-. С одной стороны, у каждой из приставок представлен набор взаимосвязанных значений, причем выбор значения в значительной мере определяется семантическими признаками глагола. С другой стороны, у каждой из приставок есть многочисленные лексикализованные употребления
Noun phrase in Enets
The paper presents a corpus-based description of the noun phrase structure in Enets dealing with both Enets dialects – Forest Enets and Tundra Enets. An Enets noun phrase has six slots for modifiers: determiner, relative clause, possessor NP, numeral, adjective phrase, apposed NP. Determiners, relative clauses, and adjective phrases are subject to linear recursion, other modifiers are not. All modifiers precede the head NP. In Enets, there is no agreement between head noun and modifiers, but numerals have different patterns in the choice of head noun number form.
Kokkuvõte. Andrej Šluinski: Noomenifraas eenetsi keeles. Artikkel esitab korpuspõhise kirjelduse eenetsi keele noomenifraasi struktuurist mõlemas eenetsi keele murdes – metsaeenetsi ja tundraeenetsi. Eenetsi noomenifraasil on kuus täiendikohta: määratleja, relatiivlause, omajat väljendav NP, numeraal, omadussõnafraas, appositsiooniline NP. Määratlejad, relatiivlaused ja omadussõnafraasid alluvad lineaarsele rekursioonile, teised täiendid mitte. Kõik täiendid eelnevad põhisõnale. Eenetsi keeles puudub põhisõna ja täiendi ühilduvus, kui numeraalid nõuavad noomenifraasi põhisõnalt erinevaid arvuvorme.
Аннотация. Андрей Шлуинский: Именная группа в энецком языке. В статье представлено выполненное на материале корпуса текстов описание структуры именной группы в обоих диалектах энецкого языка – лесном тундровом. Энецкая именная группа содержит шесть позиций для модификаторов вершинного существительного: детерминатор, относительное предложение, именная группа посессора, числительное, группа прилагательного, соположенная именная группа. Детерминаторы, относительные предложения и группы прилагательного подлежат линейной рекурсии, в отличие от других модификаторов. Все модификаторы предшествуют вершинному существительному. В энецком языке отсутствует согласование между вершинным существительным и модификаторами, но представлены разные модели выбора числовой формы вершинного существительного в именных группах с числительными
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
