1,720,973 research outputs found
Corpus linguistics for low-density varieties. Minority languages and corpus-based morphological investigations
Corpus linguistics grew up in the domain of written (and literary) varieties, while its recent methodological revolution is due to the computer-assisted capacity of elaborating massive amounts of text data. On the other hand, the so-called ‘low-density varieties’, including spoken varieties as well as varieties spoken in minority communities, have been confined to a rather marginal role. Among others, this is due to the technical problems connected to the scarce degree of normalization in linguistic –including graphemic– terms, as well as to the scarcity of language resources for automatic processing. In this paper, we will exploit the possibilities opened by corpus linguistics for acquiring and analyzing the textual patrimony of the Walser German communities of Piedmont and Aosta Valley. The varieties of Highest Alemannic spoken there, dramatically exposed to language decay, provide a limited but significant amount of data, which is accompanied by a substantial lexical documentation due to the active collaboration of the speakers’ communities in collecting and compiling local dictionaries. After briefly introducing our archive and discussing the peculiar solutions adopted for the construction of the platform, we will also present corpus-based morphological investigations regarding the representation of verbal prefixes, of the clitic group, as well as of the inflectional behaviour of verb classes
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Razvoj leksika i strukturne složenosti kod usvajanja engleskog kao prvog jezika : dvoprijelazne, kauzativne i relativne konstrukcije
This doctoral dissertation investigates various developmental parameters in the context of first language acquisition through a comparative triconstructional analysis (ditransitive, causative, and relative constructions) in children whose first language is English. The research is based on corpus analysis, during which the mentioned constructions were extracted and analysed from the English part of the CHILDES corpora. For a comprehensive interpretation and understanding of first language acquisition, intra-linguistic evidence (i.e., language phenomena related to syntax and lexicon) is observed and compared across three age-defined groups (0-3, 4-6, and 18+). Given the specificities of the three observed constructions, the analyses are tailored to their characteristics (such as frequency, complexity, and variability of certain lexemes and structural patterns appearing in these constructions). Simultaneously, they are conducted to allow not only inter-age group comparisons but also inter-constructional comparisons (such as the degree of similarity in the representation and frequency rankings of certain lexical units, as well as measures like structural and lexical complexity, including the mean length of utterance, lexical density, and type-token ratios). The main contribution of this research is the methodologically unique way of monitoring the parallel development of children's speech at the lexical, morphological, and syntactic levels, which analytically and interpretatively unifies the mentioned constructions through the lens of cognitive theories of language acquisition that are primarily based on the ‘usage-based’ model. Among other things, the research results show what can be described as a linguistic version of Pareto's principle, where all three age groups in the production of the observed constructions largely rely on a few lexical units in certain construction slots, and where the majority of their linguistic expression when it comes to said constructions is occupied by a relatively small number of structural patterns. Significant lexical and structural similarities among the same constructions were confirmed at several levels across the observed age groups, along with a more “conservative” language use and a more pronounced concentration of certain items and syntactic patterns with the declining of age. Furthermore, the ‘item-based’ claims about “simpler” constructions being more lexically restricted production-wise than more “complex” constructions was not confirmed; rather, an examination of certain lexical items used in specific construction slots showed that item-based tendencies are present in all constructions across all three age groups. In conclusion, the data on frequency and lexical and syntactic complexity obtained through comparative analysis of child and adult speech contribute to existing body of research on the nature of language acquisition and development.U ovom doktorskom radu se istražuju različiti jezično-razvojni parametri u kontekstu usvajanja prvog jezika u vidu usporedne trokonstrukcijske analize (dvoprijelazne, kauzativne i relativne konstrukcije) kod djece kojoj je prvi jezik engleski. Istraživanje se temelji na korpusnom istraživanju prilikom koje su se ekstrahirale i analizirale spomenute konstrukcije iz engleskog dijela korpusa CHILDES. U svrhu cjelovitog tumačenja i razumijevanja usvajanja prvog jezika promatraju se i uspoređuju unutarjezični dokazi (odnosno jezične pojave vezane za sintaksu i leksik) između tri dobno određene grupe (0-3, 4-6 i 18+). S obzirom na posebnosti triju promatranih konstrukcija, analize su prilagođene njihovim osobitostima (poput učestalosti, složenosti i varijabilnosti određenih leksema i strukturnih obrazaca koji se pojavljuju u spomenutim konstrukcijama), ali istovremeno provedene na način da, osim međudobnih usporedbi, dopuštaju i međukonstrukcijske usporedbe (poput stupnja različitosti u zastupljenosti i hijerarhiji učestalosti određenih leksičkih jedinica, te mjera poput strukturne i leksičke složenosti, uključujući prosječan broj morfema na sto rečenica, leksičku gustoću, te omjer različnica i pojavnica). Glavni doprinos ovog istraživanja je metodološki jedinstveno praćenje paralelnog razvoja dječjeg govora na leksičkoj, morfološkoj i sintaktičkoj razini, koje analitički i interpretativno objedinjuje navedene konstrukcije kroz prizmu kognitivističkih teorija o usvajanju jezika temeljenima prvenstveno na uporabi (eng. usage-based model). Između ostalog, rezultati istraživanja pokazuju ono što se može opisati kao lingvistička inačica Paretovog pravila, gdje se sve tri dobne skupine u proizvodnji promatranih konstrukcija uglavnom oslanjaju na nekolicinu leksičkih jedinica u određenim konstrukcijskim položajima, te da najveći dio njihovog jezičnog izričaja kod proizvodnje istih zauzima relativno mali broj strukturnih obrazaca. Na nekoliko razina je potvrđena značajna leksička i strukturna sličnost kod istih konstrukcija među dobno različitim skupinama, ali i „konzervativnija“ upotreba jezika te izraženija koncentracija određenih elemenata i sintaktičkih uzoraka s padom dobi. Osim toga, hipoteza da će „jednostavnije“ konstrukcije biti orijentirane na manji broj leksičkih elemenata (eng. item-based) od složenijih konstrukcija nije potvrđena; odnosno, pregled određenih leksičkih elemenata korištenih na određenim konstrukcijskim položajima pokazao je da su item-based tendencije prisutne u svim konstrukcijama i dobno različitim skupinama. U konačnici, podatci o učestalosti i leksičko-sintaktičkoj složenosti dobiveni komparativnom analizom govora djece i odraslih doprinose postojećim raspravama o prirodi jezičnog usvajanja i razvoja
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
Corpus analysis of adverbial multiword expressions with constituent ruka or noga
Lingvistiku kao jezičnu znanost zanima upotreba dijelova tijela u svakodnevnom govoru. Fokus ovog rada višerječni su izrazi koji uključuju dijelove tijela, a istraživanje je temeljeno na korpusu. Tri su kriterija korištena za odabir adekvatnih kandidata za istraživanje. Oni moraju:
1) biti višerječni izrazi
2) imati funkciju priložne oznake
3) imati imenicu ruka ili noga kao jedan od konstituenata.
Rad je podijeljen na dva dijela. U prvom dijelu predstavljena je teorijska podloga, i definirani su izrazi 'višerječni izraz' i 'priložna oznaka'. Višerječni izraz lingvistički je izraz koji je sastavljen od dvije ili više riječi, ali tretira se kao jedna riječ, odnosno leksička jedinica. Priložne oznake daju dodatne (neesencijalne) informacije o odnosima govornika i onome što je izrečeno. Osim dodataka, priložne oznake mogu biti i dopune (esencijalni dio) te tako imati ulogu glavnog rečeničnog dijela. Drugi dio rada svodi se na korpusno istraživanje adekvatnih izraza. Izrazi su izvađeni iz dva izvora: Baza frazema hrvatskog jezika te Bibliografija hrvatske frazeologije i frazeobibliografski rječnik. Nakon testiranja za dokazivanje višerječnosti izraza (testovi supstitucije, modifikacije i separacije), određena su značenja izraza koji su prošli testove, određena im je vrsta priložne oznake, čestoća pojavljivanja u korpusu te su navedeni najčešći glagoli, imenice, a ponekad i ostale vrste riječi, ovisno o izrazu uz koji se pojavljuju. Na kraju analize svakog izraza nalaze se primjeri koji potvrđuju navedene hipoteze. Izrazi su podijeljeni na dva dijela: izrazi s konstituentom ruka i izrazi s konstituentom noga. Oni su, pak podijeljeni po prijedlozima koji su dio izraza, odnosno manjku istih. Najviše višerječnih izraza ima funkciju priložne oznake načina (30 izraza), zatim priložne oznake mjesta (20 izraza), dva su izraza priložne oznake vremena i jedan je izraz priložna oznaka količine. Nadalje, s konstituentom postoji ruka 29 izraza (što čini 50194 pojavnice), a s konstituentom noga 24 izraza (što čini 12538 pojavnica).Linguistics as the study of language is interested in the use of the lexicon of body parts in everyday language. The focus of this paper are multi-word expressions involving body parts and is a corpus-based research. Three criteria have been used for selecting adequate candidates for the research. They must:
1) be multiword expressions
2) be sentence adverbials
3) have nouns ruka (arm, hand) or noga (leg) as one of the constituents.
The paper is divided into two parts. In the first one, the theoretical background is set out and the terms ‘multiword expression’ and ‘sentence adverbial’ are defined. A multiword expression is an expression composed of two or more words, but it is treated as one lexical unit. Sentence adverbials provide additional (non-essential) information about the relations between the speaker and the spoken. Besides the additional information, sentence adverbials can be an argument (essential part) and be the main part of the sentence. The second part of the paper is about corpus research of adequate expressions. Expressions are extracted from two sources: Baza frazema hrvatskog jezika and Bibliografija hrvatske frazeologije i frazeobibliografski rječnik. After multiword expression tests (substitution, modification, and separation), the meaning of expressions which passed the tests are defined, the types of sentence adverbial are specified, frequency of occurrence in the corpus is presented, and the most common verbs, nouns, and some other parts of speech are listed. At the end of the analysis of every expression, there are examples that confirm stated hypothesis. Expressions are divided into two parts: expressions with the constituent ruka and expressions with the constituent noga. They are divided based on prepositions, which are part of the expression, or the lack of them. The most multiword expressions are sentence adverbials of manner (30 expressions), then adverbials of place (20 expressions), two are adverbials of time, and one is an adverbial of quantity. Moreover, there are 29 expressions with the constituent ruka (corresponding to 50194 occurrences), and there are 24 expressions with constituent noga (corresponding to 12538 occurrences)
- …
