Charles University

Biblio at Institute of Formal and Applied Linguistics
Not a member yet
    539 research outputs found

    Adaptation of machine translation for multilingual information retrieval in medical domain

    No full text
    In this work, we investigate machine translation (MT) of search queries in the context of cross-lingual information retrieval (IR) in the domain of medicine. The main focus is on MT adaptation techniques to increase translation quality, however we also explore MT adaptation to improve cross-lingual IR directly. The experiments described herein have been performed and thoroughly evaluated for MT quality on the datasets created within the Khresmoi project and for IR performance on the CLEF eHealth 2013 datasets on three language pairs: Czech–English, German–English, and French–English. The search query translation results achieved in our experiments are outstanding – our systems outperformed not only our strong baselines, but also the Google Translate and Microsoft Bing Translator in direct comparison carried out on all the language pairs. In terms of the retrieval performance on this particular test collection, a significant improvement over the baseline has been achieved only for French–English. Throughout the article, we provide discussion and details on the contribution of the state-of-the-art features and adaptation techniques under exploration and provide future research directions

    Machine Translation within One Language as a Paraphrasing Technique

    No full text
    We present a method for improving machine translation (MT) evaluation by targeted paraphrasing of reference sentences. For this purpose, we employ MT systems themselves and adapt them for translating within a single language. We describe this attempt on two types of MT systems -- phrase-based and rule-based. Initially, we experiment with the freely available SMT system Moses. We create translation models from two available sources of Czech paraphrases -- Czech WordNet and the Meteor Paraphrase tables. We extended Moses by a new feature that makes the translation targeted. However, the results of this method are inconclusive. In the view of errors appearing in the new paraphrased sentences, we propose another solution -- targeted paraphrasing using parts of a rule-based translation system included in the NLP framework Treex

    Khresmoi Summary Translation Test Data 1.1

    No full text
    This package contains data sets for development and testing of machine translation of sentences from summaries of medical articles between Czech, English, French, and German

    Observations and Lessons Learnt from Non Health Professionals Evaluating a Health Search Engine

    No full text
    This article presents the results of one of the stages of the user-centered evaluation conducted in a framework of the EU project Khresmoi. In a controlled environment, users were asked to perform health-related tasks using a search engine specifically developed for trustworthy online health information. Twenty seven participants from largely the Czech Republic and France took part in the evaluation. All reported overall a positive experience, while some features caused some criticism. Learning points are summed up regarding running such types of evaluations with the general public and specifically with patients

    HamleDT 2.0

    No full text
    HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently. We use the newest basic Universal Stanford Dependencies, without added language-specific subtypes

    HamleDT: Harmonized Multi-Language Dependency Treebank

    No full text
    We present HamleDT – a HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. In the present article, we provide a thorough investigation and discussion of a number of phenomena that are comparable across languages, though their annotation in treebanks often differs. We claim that transformation procedures can be designed to automatically identify most such phenomena and convert them to a unified annotation style. This unification is beneficial both to comparative corpus linguistics and to machine learning of syntactic parsing

    CUNI in WMT14: Chimera Still Awaits Bellerophon

    Get PDF
    We present our English→Czech and English→Hindi submissions for this year’s WMT translation task. For English→Czech, we build upon last year’s CHIMERA and evaluate several setups. English→Hindi is a new language pair for this year. We experimented with reverse self-training to acquire more (synthetic) parallel data and with modeling target-side morphology

    CUNI at the ShARe/CLEF eHealth Evaluation Lab 2014

    No full text
    This report describes the participation of the team of Charles University in Prague at the ShARe/CLEF eHealth Evaluation Lab in 2014

    Improving Evaluation of English-Czech MT through Paraphrasing

    Get PDF
    In this paper, we present a method of improving the accuracy of machine translation evaluation of Czech sentences. Given a reference sentence, our algorithm transforms it by targeted paraphrasing into a new synthetic reference sentence that is closer in wording to the machine translation output, but at the same time preserves the meaning of the original reference sentence. Grammatical correctness of~the new reference sentence is provided by applying Depfix on newly created paraphrases. Depfix is a system for post-editing English-to-Czech machine translation outputs. We adjusted it to fix the errors in paraphrased sentences. Due to a noisy source of our paraphrases, we experiment with adding word alignment. However, the alignment reduces the number of paraphrases found and the best results were achieved by~a~simple greedy method with only one-word paraphrases thanks to their intensive filtering. BLEU scores computed using these new reference sentences show significantly higher correlation with human judgment than scores computed on the original reference sentences

    Twitter Crowd Translation -- Design and Objectives

    No full text
    The paper describes the design and implementation of a system for human and machine translation of tweets. In these early experiments, we limit the system to follow some selected sources from Ukraine (the source language is primarily Ukrainian and Russian, sometimes English)

    58

    full texts

    539

    metadata records
    Updated in last 30 days.
    Biblio at Institute of Formal and Applied Linguistics
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇