1,720,960 research outputs found

    Ortografia-erroreak eta konpetentzia-erroreak Webeko euskarazko testuetan

    Get PDF
    Lan honetan euskarazko ortografia-erroreen azterketa egin dugu webetik jasotako dokumentuekin osatutako hainbat corpusetan (testu-bildumetan), eta horrela corpus horien kalitatea estimatu dugu. Metodologia finkatzeko, ingeleserako eta alemanerako egin den antzeko lanean oinarritu gara (Ringlstetter et al., 2006), baina, euskararen ezaugarriak direla eta, ez dugu teknologia bera erabili erroreak identifikatzeko. Euskarak morfologia aberatsa duenez, erroreak identifikatzeko berrerabili egin ditugu aurretik garatutako ortografia-zuzentzaileak. Bide horretatik, detekzioaren estaldura handiagoa da eta, gainera, prozesuaren garapena azkarragoa izan da berrerabilpena dela-eta. Horrekin batera, posible da ia automatikoki halako tresnak dituzten beste hizkuntzetan metodo bera erabiltzea. Analisiaren emaitzak balio dezake zuzentasunaren araberako testuen sailkapena egiteko, eta bide batez, aukera ematen du gutxieneko kalitate bat ez duten testuak baztertzeko

    Ortografia-erroreak eta konpetentzia-erroreak Webeko euskarazko testuetan

    Get PDF
    Lan honetan euskarazko ortografia-erroreen azterketa egin dugu webetik jasotako dokumentuekin osatutako hainbat corpusetan (testu-bildumetan), eta horrela corpus horien kalitatea estimatu dugu. Metodologia finkatzeko, ingeleserako eta alemanerako egin den antzeko lanean oinarritu gara (Ringlstetter et al., 2006), baina, euskararen ezaugarriak direla eta, ez dugu teknologia bera erabili erroreak identifikatzeko. Euskarak morfologia aberatsa duenez, erroreak identifikatzeko berrerabili egin ditugu aurretik garatutako ortografia-zuzentzaileak. Bide horretatik, detekzioaren estaldura handiagoa da eta, gainera, prozesuaren garapena azkarragoa izan da berrerabilpena dela-eta. Horrekin batera, posible da ia automatikoki halako tresnak dituzten beste hizkuntzetan metodo bera erabiltzea. Analisiaren emaitzak balio dezake zuzentasunaren araberako testuen sailkapena egiteko, eta bide batez, aukera ematen du gutxieneko kalitate bat ez duten testuak baztertzeko

    Weba euskarazko corpus gisa

    Get PDF
    The Basque language. just as any other, needs text corpora to survive in the modern world and to be used normally. But Basque corpora are few and small compared to those in other major languages. This is so because other languages have made use of the "Web-as-Corpus" approach , which consists of using the web as a corpus or as a source of texts for corpora. ln this paper, we describe the research carried out in his PhD thesis by the first author, under the supervision of the other two authors, to use the web and automatic methods for Basque corpus building, and also the tools developed and the results obtained. Out of them we can conclude that the "Web-as-Corpus" approach is val id to improve the state of Basque corpora , since with the developed tools we have collected quality corpora of different types (very large general corpora, specialized corpora, comparable corpora ... ) and built a service to query the web as a Basque corpus.Many of these tools and services ha ve already been placed online for their public use.; Euskarak, beste edozein hizkuntzak bezala , testu-corpusak behar ditu mundu modernoan bizirauteko eta normalki erabiltzeko. Alabaina , euskarazko corpusak gutxi eta txikiak dira , beste hizkuntza handiagoenekin konparatuz gero. Hori horrela da beste hizkuntzek "Web-as-Corpus" izeneko planteamendua baliatu dutelako, hau da, weba erabili dutelako corpus gisa edo corpusak osatzeko testu-iturritzat . Artikulu honetan azaltzen dira bere doktorego-tesian lehenengo autoreak, beste bi autoreen zuzendaritzapean, euskarazko corpusgintzarako weba eta metodo automatikoak baliatzeko egindako ikerketak, aratutako tresnak eta lortutako emaitzak . Horietatik ondorioztatu daiteke "Web-as-Corpus" planteamendua baliagarria dela euskarazko corpusen egoera hobetzeko, garatu diren tresna informatikoen bidez weba corpus gisa kontsultatzeko tresna bat eraiki baita eta mota askotako eta kalitatezko corpusak lortu ahal izan baitira (corpus orokor oso handiak, corpus espezializatuak, corpus konparagarriak, .. ). Horietako asko jada online gizartearen eskura jarri dira

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Genero aldetik anbiguoa den hizketaren sintesia euskaraz hizlari-bektore manipulazioaren bidez

    No full text
    There is a growing interest in text-to-speech (TTS) systems with gender-ambiguous voices, among other things due to their potential to avoid gender biases and stereotypes in voice assistants and smart speakers. In this paper we present and evaluate some novel methods that apply voice morphing techniques to speaker embeddings in order to obtain neural network-based gender-ambiguous voiced TTS systems for the Basque language. The speaker embeddings are obtained training a multi-speaker Tacotron 2. We compare the performance of systems with and without speaker embedding normalization with a scaling parameter, and also the application of these systems to the average embeddings of each gender and to real voice embeddings. The results prove that the methods presented are valid to obtain gender-ambiguous voices with acceptable, albeit improvable, quality.; Genero aldetik anbiguoa den ahotsa duten text-to-speech (TTS) sistemek gero eta interes handiagoa pizten dute; besteak beste, laguntzaile birtualetan eta bozgorailu adimendunetan genero-alborapenak eta estereotipoak saihesteko duten ahalmenagatik. Artikulu honetan, ahots-bihurketarako teknika berriak aplikatu dizkiegu ahots-bektoreei, sare neuronaletan oinarrituta dauden eta genero aldetik anbiguoak diren euskarazko TTS sistemak lortzeko. Hizlari-bektoreak hiztun anitzeko Tacotron 2-a entrenatuz lortu ditugu. Hizlari-bektoreen normalizazioa eta eskala-parametro bat erabiltzen duten eta erabiltzen ez duten sistemak konparatu ditugu, baita genero bakoitzeko batez besteko hizlari bektore eta ahots errealen hizlari bektoreen erabilera sistema horietan. Emaitzek frogatzen dute aurkeztutako metodoak baliozkoak direla genero aldetik anbiguoak diren ahotsak lortzeko eta kalitate onargarria dutela baina hobetu daitezkeela

    Orthographic and competence errors in the Basque Web

    No full text
    En este trabajo se estima la calidad de los corpus en euskera obtenidos de la Web siguiendo una metodología similar a la propuesta por Ringlstetter et al. (2006) para el inglés y el alemán. Sin embargo nuestro trabajo difiere del mencionado en que al tratar un idioma de gran riqueza morfológica hemos optado por reutilizar verificadores ortográficos para reconocer los errores. Esto trae consigo, en nuestra opinión, una cobertura mayor de los errores que se estudian, además de la reutilización de recursos previamente desarrollados, lo que hace el método interesante para aplicarlo, sin prácticamente trabajo manual, a lenguas que tienen disponibles estos recursos. Los resultados van a ser de gran interés para detectar los distintos tipos de textos obtenidos de la Web en euskera según su corrección, y filtrar aquellos que pueden generar problemas o no tienen una calidad mínima.The objective of the work presented in this paper is to estimate the quality of corpora retrieved from the Basque Web. The methodology followed is similar to that used for English and Germany by Ringlstetter et al. (2006). The main difference lies in the fact that we reuse spelling checkers for detecting errors. We think that by this way we obtain a higher error coverage and that the method can be applied to other languages with practically no manual work provided such tools are available for them. The results obtained can be useful for improving the quality of corpora obtained from the web, eliminating documents containing errors over a given threshold.Proyecto parcialmente subvencionado por los proyectos OpenMT2 (Ministerio de Ciencia e Innovación, TIN2009-14675-C03-01) y Berbatek (Eusko Jaurlaritza, IE09-262)

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
    corecore