Charles University

Biblio at Institute of Formal and Applied Linguistics
Not a member yet
    539 research outputs found

    Srovnání koreferenčních výrazů v češtině a angličtině

    No full text
    In this work, we present a comprehensive study on correspondences between certain classes of coreferential expressions in English and Czech. We focus on central pronouns, relative pronouns, and anaphoric zeros. We designed an alignment-refining algorithm for English personal and possessive pronouns and Czech relative pronouns that improves the quality of alignment links not only for the classes it aimed at but also in general. Moreover, the instances of anaphoric expressions we focus on were manually annotated with their alignment counterparts, which served as a basis for this empirical study. The collected statistics of correspondences are contrasted with theoretical assumptions regarding the use of anaphoric means in the languages under analysis, such as pro-drop properties, the use of finite and non-finite constructions, etc. Finally, we present the ways how the observed correspondences can be exploited in cross-lingual coreference resolution

    KLcpos3 - a Language Similarity Measure for Delexicalized Parser Transfer

    No full text
    We present KLcpos3, a language similarity measure based on Kullback-Leibler divergence of coarse part-of-speech tag trigram distributions in tagged corpora. It has been designed for multilingual delexicalized parsing, both for source treebank selection in single-source parser transfer, and for source treebank weighting in multi-source transfer. In the selection task, KLcpos3 identifies the best source treebank in 8 out of 18 cases. In the weighting task, it brings +4.5% UAS absolute, compared to unweighted parse tree combination

    MSTperl parser (2015-05-19)

    No full text
    MSTperl is a Perl reimplementation of the MST parser of Ryan McDonald, with several additional advanced functions, such as support for parallel features

    Comparison of Coreference Resolvers for Deep Syntax Translation

    No full text
    This work focuses on using anaphora for machine translation with deep-syntactic transfer. We compare multiple coreference resolvers for English in terms of how they affect the quality of pronoun translation in English-Czech and English-Dutch machine translation systems with deep transfer. We examine which pronouns in the target language depend on anaphoric information, and design rules that take advantage of this information. The resolvers’ performance measured by translation quality is contrasted with their intrinsic evaluation results. In addition, a more detailed manual analysis of English-to-Czech translation was carried out

    Coreference chains in Czech, English and Russian: Preliminary findings

    Get PDF
    Tento článek je pilotní srovnavací výzkum koreferenčních řetězců v češtině, angličtině a ruštině. Podrobili jsme analýze 16 srovnatelných textů ve třech jazycích. Naší motivací bylo zjistit lingvistickou strukturu koreferenčních řetězců v těchto jazycích a určit, které faktory ovlivňují tuto strukturu

    The LINDAT/CLARIN large infrastructure: data and services

    No full text
    LINDAT/CLARIN, as a node of the pan-European research infrastructure Clarin ERIC, has been presented. Its repository has been featured together with data archivation techniques and related web services and applications

    WMT2013 Test Set in Vietnamese

    No full text
    The file contains the standard WMT 2013 test set, manually translated from English to Vietnamese, thus extending the WMT 2013 multi-parallel set of languages (English, Czech, German, French, Spanish, Russian)

    Deep Linguistic Information in Machine Translation

    No full text
    Fundamentals of Machine Translation technology using Deep language analysis using the traditional "Vauquois triangle" have been presented, which now use advanced statistical and machine learning techniques. In addition, latest WMT 2015 Shared Task competition results for the en-cs pair have been shown in which the Chimera system created at UFAL MFF UK won

    Retrieving information about medical symptoms

    No full text
    This paper details methods, results and analysis of the CLEF 2015 eHealth Evaluation Lab, Task 2. This task investigates the effectiveness of web search engines in providing access to medical information with the aim of fostering advances in the development of these technologie

    Abstract Meaning Representation: comparing languages and their representations

    No full text
    Abstract Meaning Representation is a newly developed formalism for representing meaning, which abstracts from syntax and some other phenomena but is (still) language-dependent. It has been developed by a consortium of mostly U.S. universities, with team members including Martha Palmer, Kevin Knight, Philipp Koehn, Ulf Hermjakob, Kathy McKeown, Nianwen Xue and others. In the talk, the basic facts about the AMR will be presented, and then comparison will be made for Czech and English as carried out in detail on a small 100-sentence corpus; some examples from Chinese-English comparison will also be shown. In addition, AMR will be compared to the deep syntactic representation used in the set of Prague Dependency Treebanks, and observations will be made about the level of abstraction used in these two formalisms. There is an ongoing work on possible conversion between the Prague deep syntactic representation and the AMR representation, and the main issues of such a conversion will also be described, as well as plans for future studies and possible corpus annotation work in the nearest future

    58

    full texts

    539

    metadata records
    Updated in last 30 days.
    Biblio at Institute of Formal and Applied Linguistics
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇