1,721,179 research outputs found
Kurdish Studies
Kurdish Studies has no intention to regularly cover and comment on recent events. However, we are definitely interested in publishing studies, based on serious research and critical reflection, that provide important background or new insights relevant to understanding these events. We would specifically encourage colleagues who could contribute to deepening our understanding of the developments in Syria (or, for that matter, developments affecting the Kurds of Iran, who rarely if ever hit the headlines and who are the most seriously under-studied part of Kurdish society).
This special issue of Kurdish Studies is dedicated to studies of the Kurdish language, the oldest branch of Kurdish studies, and the first to find a degree of academic institutionalisation. Compared to other major Middle Eastern languages, Kurdish has received relatively little serious investigation, but there is a gradually growing corpus of empirical and theoretical research, of which the guest editors give a useful overview in the introduction
Indigenous narrative texts in linguistic typology
Recent years have witnessed the developments of a new strand in linguistic typology called “text-based” or “corpus-based” typology (Authors, 2011). The central idea is to supplement grammar-based language typology (cf Dryer & Haspelmath 2013) by the systematic analysis of text corpora from diverse languages (Waelchli 2006, 2009; Cysouw & Waelchli 2007). Corpus- or text- based typological studies can reveal differences and universals of language use that are not usually captured in descriptive grammars; they thus make important contributions to usage-based accounts of linguistic diversity and universals. Text-based typology has important precursors in the studies of discourse structure and referentiality pioneered by Wallace Chafe and Talmy Givón in the 1970ies. Despite the immense ramification for usage-based accounts of discourse and grammatical structures, research in the Chafe-Givón tradition has mostly been restricted to relatively small text-data sets, often from relatively well-studied languages of Europe or East Asia. Furthermore, to date there has been comparatively little attempt to harness more recent technical developments in documentary and corpus linguistics to the cross-linguistic investigation of natural discourse. We here report on a text-based typological initiative that taps into the spoken language data compiled in unprecedented amounts and quality within various language documentation projects. In contrast to other efforts in text-based typology that use parallel (Waelchli 2009; Cysouw & Waelchli 2007) or stimuli- based texts (Pear Film re-tellings; Bickel 2003; Du Bois 1987; or Frog Story re- tellings; Berman & Slobin 1994), we aggregate spoken language corpora of what we call “original texts”, that is texts that resemble indigenous literature and have been recorded in their respective cultural settings during months- or years-long fieldwork. We demonstrate the viability of applying unified system of morpho- syntactic text annotation to texts from diverse languages, thus enabling quantitative research in the Chafe-Givón tradition in much more systematic ways, and on a much broader text data base that reflects more reliably the actual use of diverse languages by their speakers. We demonstrate the impact of our approach through examples of studies that are published in high-quality journals and other publications; time permitting, we will also illustrate further on-going research. Our paper shows that spoken texts from language documentation are not only a viable resource for language maintenance, and serving as a basis for grammars and dictionaries, but that also, assuming certain conditions of technical implementation and annotation are observed, bear unique potentials as direct input for language typology and theoretical linguistics. References: BERMAN, RUTH & DAN SLOBIN (eds.) 1994. (eds), Relating events in narrative: a crosslinguistic developmental study. Hillsdale, NJ: Erlbaum. BICKEL, BALTHASAR. 2003. Referential density in discourse and syntactic typology. Language 79.708–36. DOI: 10.1353/lan.2003.0205. CYSOUW, MICHAEL, and BERNHARD WÄLCHLI. 2007. Parallel texts: Using translational esquivalents in linguistic typology. STUF - Sprachtypologie und Universalienforschung 60.2.95–99. DOI: 10.1524/ stuf.2007.60.2.95. DU BOIS, JOHN W. 1987. The discourse basis of ergativity. Language 63.4, 805-855. AUTHORS. 2011. Comparing corpora from endangered languages: Explorations in language typology based on original texts. In AUTHORS (eds.), 55-86. AUTHORS (eds.) 2011. Documenting endangered languages: Achievements and perspectives. Berlin: Mouton de Gruyter DRYER, MATHEW & MARTIN HASPELMATH (eds.). 2013. The World Atlas of Language Structures Online. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at http://wals.info, Accessed on 2016-09-01.) WÄLCHLI, BERNHARD. 2006. Descriptive typology, or, the typologist’s expanded toolkit. Konstanz: University of Konstanz, MS. WÄLCHLI, BERNHARD. 2009. Motion events in parallel texts: A study in primary data typology. Bern: Uni-versity of Bern Habilitation thesis
Multi-CAST Northern Kurdish (audio recordings)
<p>This archive contains audio recordings for the <strong>Multi-CAST Northern Kurdish</strong> corpus (Haig, Vollmer, & Thiele 2019), originally published in May 2015 with version 1505 of the <i>Multi-CAST</i> collection (Haig & Schnell 2015). The annotation and documentation files accompanying these files have been archived separately. The recordings are available as WAV and MP3 files. </p><p><strong>Northern Kurdish</strong> [<a href="https://glottolog.org/resource/languoid/id/nort2641">nort2641</a>], also known as Kurmanjî, is a Northwest Iranian language spoken in eastern Turkey, Iraq, Syria, and parts of western Iran. The three texts recorded here are traditional narratives, from a female and a male speaker who grew up near the townships of Erzurum and Muš, respectively. The texts were recorded in Germany in the late 1990s and early 2000s, and subsequently transcribed, translated, and annotated for Multi-CAST by Geoffrey Haig, Abdullah Incekan, Hanna Thiele, and Maria Vollmer. A description of the language can be found in Haig (2018).</p><p> </p><p><strong>Citation</strong></p><ul><li>Haig, Geoffrey & Vollmer, Maria & Thiele, Hanna. 2019. Multi-CAST Northern Kurdish. In Haig, Geoffrey & Schnell, Stefan (eds.), <i>Multi-CAST: Multilingual corpus of annotated spoken texts.</i> [version of the annotations used]. Bamberg: University of Bamberg.</li></ul><p><strong>References</strong></p><ul><li>Haig, Geoffrey. 2018. <i>Northern Kurdish (Kurmanjî)</i>. In Haig, Geoffrey & Khan, Geoffrey (eds.), <i>The languages and linguistics of Western Asia: An areal perspective</i>, 106–158. Berlin: Mouton de Gruyter.</li><li>Haig, Geoffrey & Schnell, Stefan (eds.). 2015. <i>Multi-CAST: Multilingual corpus of annotated spoken texts.</i> [version]. Bamberg: University of Bamberg.</li></ul><p> </p>
Sprachtypologie und Universalienforschung/Language typology and language universals, Special edition on Kurdish Linguistics
Multi-CAST: Multilingual corpus of annotated spoken texts
<p><strong>Multi-CAST</strong>, the <em>Multilingual Corpus of Annotated Spoken Texts</em> (Haig & Schnell 2015), is a collection of annotated spoken-language corpora from a typologically diverse set of languages. Most of the data stem from documentation projects undertaken on lesser-researched and endangered languages. The texts are overwhelmingly unscripted, non-elicited, monologic narratives. </p>
<p>Each corpus in the collection is an individually citable resource that was contributed by experts on the respective languages in cooperation with the collection editors. The Multi-CAST collection as a whole was designed and compiled by Geoffrey Haig and Stefan Schnell with the assistance of Nils Schiborr, and is to date the only freely-available, multilingual, spoken-language corpus that combines morphological and morphosyntactic glossing with annotation of discourse referents. Each Multi-CAST corpus includes audio recordings (as WAV and MP3 files; archived separately, see below), annotation files in a number of file formats (including as EAF files for use with the free <a href="https://archive.mpi.nl/tla/elan">linguistic annotation software ELAN</a>, and as TSV and XML files), metadata on the speakers and texts, as well as documentation on the language, speech communities, recording situations, and analytical decisions pertinent to the annotations.</p>
<p>The annotation files use a multi-tier structure built on a time-aligned segmentation of the text into utterance units, from which derive a transcription and idiomatic English translation. Utterance units are segmented further into grammatical words with morphological glossing (following the <a href="https://www.eva.mpg.de/lingua/resources/glossing-rules.php">Leipzig Glossing Rules</a>) and annotations with the GRAID (<em>Grammatical Relations and Animacy in Discourse</em>, <a href="https://nbn-resolving.org/urn:nbn:de:bvb:473-opus4-262354">Haig & Schnell 2014</a>) and RefIND (<em>Referent Indexing in Natural Language Discourse</em>, <a href="https://fis.uni-bamberg.de/handle/uniba/91175">Schiborr et al. 2018</a>) annotation schemes. Further information on the contents of the collection and the structure of the annotations can be found in the <em>Multi-CAST collection overview</em> (Schiborr 2023), which is included in this archive. The <a href="https://CRAN.R-project.org/package=multicastR"><em>multicastR</em> package</a> (Schiborr 2018) provides a simple interface for directly accessing the Multi-CAST annotation data through the <a href="https://www.r-project.org/">statistical computing language R</a>.</p>
<p><strong>This archive contains version 2311</strong> <strong>of the Multi-CAST collection</strong> (originally published in November 2023) and comprises data from 18 languages, encompassing around 19 hours of recordings, 29000 clause units, and 140000 words across 136 individual texts. The audio files accompanying these data sets have been archived separately; they can be found via the links in the list below.</p>
<ul>
<li><strong>Arta</strong> [<a href="https://glottolog.org/resource/languoid/id/arta1239">arta1239</a>] (Kimoto 2019)<br>— link to audio: <em><a href="https://doi.org/10.48564/unibafd-c0jd0-7qt52">10.48564/unibafd-c0jd0-7qt52</a></em></li>
<li><strong>Bora</strong> [<a href="https://glottolog.org/resource/languoid/id/bora1263">bora1263</a>] (Seifart & Hong 2022)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-zcyz8-x7f04">10.48564/unibafd-zcyz8-x7f04</a></li>
<li><strong>Cypriot Greek</strong> [<a href="https://glottolog.org/resource/languoid/id/cypr1249">cypr1249</a>] (Hadjidas & Vollmer 2015)<br>— <em>no audio files available</em></li>
<li><strong>English</strong> [<a href="https://glottolog.org/resource/languoid/id/sout3282">sout3282</a>] (Schiborr 2015)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-4nays-jwa80">10.48564/unibafd-4nays-jwa80</a></li>
<li><strong>Jinghpaw</strong> [<a href="https://glottolog.org/resource/languoid/id/kach1280">kach1280</a>] (Kurabe 2021)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-jav5f-paa07">10.48564/unibafd-jav5f-paa07</a></li>
<li><strong>Kalamang</strong> [<a href="https://glottolog.org/resource/languoid/id/kara1499">kara1499</a>] (Visser 2021)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-z9wt8-jwd54">10.48564/unibafd-z9wt8-jwd54</a></li>
<li><strong>Mandarin</strong> [<a href="https://glottolog.org/resource/languoid/id/mand1415">mand1415</a>] (Vollmer 2020)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-bxcvm-m9e27">10.48564/unibafd-bxcvm-m9e27</a></li>
<li><strong>Matukar Panau</strong> [<a href="https://glottolog.org/resource/languoid/id/matu1261">matu1261</a>] (Barth, Davey & Matheas 2023) <strong>[NEW!]</strong><br>— link to audio: <a href="https://doi.org/10.48564/unibafd-0sa31-g8r71">doi.org/10.48564/unibafd-0sa31-g8r71</a></li>
<li><strong>Nafsan</strong> [<a href="https://glottolog.org/resource/languoid/id/sout2856">sout2856</a>] (Thieberger & Brickell 2019)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-jq8x6-d8p78">10.48564/unibafd-jq8x6-d8p78</a></li>
<li><strong>Northern Kurdish</strong> [<a href="https://glottolog.org/resource/languoid/id/nort2641">nort2641</a>] (Haig, Vollmer & Thiele 2019)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-6sbd4-r0868">10.48564/unibafd-6sbd4-r0868</a></li>
<li><strong>Persian</strong> [<a href="https://glottolog.org/resource/languoid/id/tehr1242">tehr1242</a>] (Adibifar 2016)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-37wvv-n0j98">10.48564/unibafd-37wvv-n0j98</a></li>
<li><strong>Sanzhi Dargwa</strong> [<a href="https://glottolog.org/resource/languoid/id/sanz1248">sanz1248</a>] (Forker & Schiborr 2019)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-fahwc-1ha62">10.48564/unibafd-fahwc-1ha62</a></li>
<li><strong>Sumbawa</strong> [<a href="https://glottolog.org/resource/languoid/id/sumb1241">sumb1241</a>] (Shiohara 2022)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-p63kx-vzd97">10.48564/unibafd-p63kx-vzd97</a></li>
<li><strong>Tabasaran</strong> [<a href="https://glottolog.org/resource/languoid/id/taba1259">taba1259</a>] (Bogomolova, Ganenkov & Schiborr 2021)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-vqjky-k8g84">10.48564/unibafd-vqjky-k8g84</a></li>
<li><strong>Teop</strong> [<a href="https://glottolog.org/resource/languoid/id/teop1238">teop1238</a>] (Mosel & Schnell 2015)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-03n2z-bm579">10.48564/unibafd-03n2z-bm579</a></li>
<li><strong>Tondano</strong> [<a href="https://glottolog.org/resource/languoid/id/tond1251">tond1251</a>] (Brickell 2016)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-1nkkj-f9352">10.48564/unibafd-1nkkj-f9352</a></li>
<li><strong>Tulil</strong> [<a href="https://glottolog.org/resource/languoid/id/taul1251">taul1251</a>] (Meng 2019)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-h1wq5-wzh05">10.48564/unibafd-h1wq5-wzh05</a></li>
<li><strong>Vera'a</strong> [<a href="https://glottolog.org/resource/languoid/id/vera1241">vera1241</a>] (Schnell 2015)<br>— link to audio: <a href="https://doi.org/10.48564/unibafd-es22h-1j872">10.48564/unibafd-es22h-1j872</a></li>
</ul>
<p> </p>
<p><strong>Citation for the entire Multi-CAST collection:</strong></p>
<ul>
<li>Haig, Geoffrey & Schnell, Stefan (eds.). 2015. <em>Multi-CAST: Multilingual corpus of annotated spoken texts.</em> Version 2311. Bamberg: University of Bamberg. (DOI: <a href="https://doi.org/10.48564/unibafd-q1h0x-3kf71">10.48564/unibafd-q1h0x-3kf71</a>)<br> </li>
</ul>
<p><strong>Citations for individual Multi-CAST corpora:</strong></p>
<ul>
<li>Adibifar, Shirin. 2016. Multi-CAST Persian. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Barth, Danielle & Davey, Kira & Matheas, Maria. 2023. Multi-CAST Matukar Panau. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Bogomolova, Natalia & Ganenkov, Dmitry & Schiborr, Nils N. 2021. Multi-CAST Tabasaran. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Brickell, Timothy. 2016. Multi-CAST Tondano. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Forker, Diana & Schiborr, Nils N. 2019. Multi-CAST Sanzhi Dargwa. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Hadjidas, Harris & Vollmer, Maria. 2015. Multi-CAST Cypriot Greek. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Haig, Geoffrey & Vollmer, Maria & Thiele, Hanna. 2019. Multi-CAST Northern Kurdish. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Kimoto, Yukinori. 2019. Multi-CAST Arta. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Kurabe, Keita. 2021. Multi-CAST Jinghpaw. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Meng, Chenxi. 2019. Multi-CAST Tulil. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Mosel, Ulrike & Schnell, Stefan. 2015. Multi-CAST Teop. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Schiborr, Nils N. 2015. Multi-CAST English. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Schnell, Stefan. 2015. Multi-CAST Vera'a. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Seifart, Frank & Hong, Tai. 2022. Multi-CAST Bora. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Shiohara, Asako. 2022. Multi-CAST Sumbawa. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Thieberger, Nick & Brickell, Timothy. 2019. Multi-CAST Nafsan. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Visser, Eline. 2021. Multi-CAST Kalamang. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
<li>Vollmer, Maria. 2020. Multi-CAST Mandarin. In Haig, Geoffrey & Schnell, Stefan (eds.), <em>Multi-CAST.</em></li>
</ul>
<p> </p>
The Corpus of Contemporary Kurdish Newspaper Texts
<p>The <strong>CCKNT</strong> comprises written Northern Kurdish (Kurmanjî) journalistic texts, compiled from online newspaper texts in 1999. The corpus consists of 483 texts, totalling around 214 000 words. It contains texts from two Kurdish publications: <em>Azadiya Welat</em>, a weekly Kurdish newspaper, and CTV, a company that broadcasts news items in Kurdish on the internet. The texts are not tagged or translated.</p>
<p>The corpus was compiled as part of a project on modern Kurdish syntax, conducted from 1999–2001 at the Seminar für Allgemeine und Vergleichende Sprachwissenschaft at the University of Kiel.</p>
<p> </p>
<p><strong>Citation</strong></p>
<ul>
<li>Haig, Geoffrey. 2001. <em>The Corpus of Contemporary Kurdish Newspaper Texts (CCKNT).</em> Kiel: University of Kiel. (DOI: <a href="https://doi.org/10.48564/unibafd-v6rmw-cx940">10.48564/unibafd-v6rmw-cx940</a>) (date accessed)</li>
</ul>
<p> </p>
Comparing corpora from endangered language projects : explorations in language typology based on original texts
- …
