539 research outputs found
Sort by
Coreferential expressions in English and Czech
In this talk, we present a comprehensive study on mappings between certain classes of coreferential expressions in English and Czech. We focused on central pronouns, relative pronouns and anaphoric zeros. For instance, the English sentence "It switched to a caffeine-free formula using its new Coke in 1985" has been in PCEDT translated to "V roce 1985 přešla na bezkofeinovou recepturu, kterou používá pro svojí novou kolu". This pair of sentences exhibits several types of changes in expressing coreference: English personal pronouns turns into a Czech zero, possessive pronoun into a possessive reflexive and finally, the -ing participle has been translated to a relative clause. In a similar manner, we have collected a statistics of mappings from a subsection of PCEDT, which we will support by multiple examples and contrast with the theoretical assumptions. For such a study, the quality of word alignment is crucial. Thus, we designed a rule-based refining algorithm for English personal and possessive pronouns and Czech relative pronouns, which served as an automatic alignment pre-annotation. Subsequently, this annotation has been manually corrected and completed, obtaining a basis for this empirical study
TmTriangulate: A Tool for Phrase Table Triangulation
Over the past years, pivoting methods, i.e. machine translation via a third language, gained respectable attention. Various experiments with different approaches and datasets have been carried out but the lack of open-source tools makes it difficult to replicate the results of these experiments. This paper presents a new tool for pivoting for phrase-based statistical machine translation by so called phrase-table triangulation. Besides the tool description, this paper discusses the strong and weak points of various triangulation techniques implemented in the tool
MSTParser Model Interpolation for Multi-source Delexicalized Transfer
We introduce interpolation of trained MSTParser models as a resource combination method for multi-source delexicalized parser transfer. We present both an unweighted method, as well as a variant in which each source model is weighted by the similarity of the source language to the target language. Evaluation on the HamleDT treebank collection shows that theweightedmodelinterpolationperforms comparably to weighted parse tree combination method, while being computationally much less demanding
Using Parallel Texts and Lexicons for Verbal Word Sense Disambiguation
We present a system for verbal Word Sense Disambiguation (WSD) that is able to exploit additional information from parallel texts and lexicons. It is an extension of our previous WSD method, which gave promising results but used only monolingual features. In the follow-up work described here, we have explored two additional ideas: using English-Czech bilingual resources (as features only - the task itself remains a monolingual WSD task), and using a 'hybrid' approach, adding features extracted both from a parallel corpus and from manually aligned bilingual valency lexicon entries, which contain subcategorization information. Albeit not all types of features proved useful, both ideas and additions have led to significant improvements for both languages explored
CsEnVi Pairwise Parallel Corpora
CsEnVi Pairwise Parallel Corpora is a collection of two parallel corpora: an English-Vietnamese one and a Czech-Vietnamese one. The corpora contain translations of movie and TED subtitles that were already available. The corpora were cleaned using our semi-automatic filtering and are provided aligned at the sentence level
A Summary of Research Activities: Technologies – Demands – Gaps – Roadmaps
A summary of research activities in the area of machine translation in the EU (Technologies – Demands – Gaps – Roadmaps) has been presented, including contributions from multiple other EU-funded projects
Giving a Sense: A Pilot Study in Concept Annotation from Multiple Resources
We present a pilot study of a web-based annotation of words with senses. The annotated senses come from several knowledge bases and sense inventories. The study is the first step in a planned larger annotation of grounding and should allow us to select a subset of the sense sources that cover any given text reasonably well and show an acceptable level of inter-annotator agreement
CUNI at the CLEF eHealth 2015 Task 2
This report describes the participation of the team of Charles University in Prague at the CLEF eHealth 2015 Task 2
MT-ComparEval: Graphical evaluation interface for Machine Translation development
The tool described in this article has been designed to help MT developers by implementing
a web-based graphical user interface that allows to systematically compare and evaluate various
MT engines/experiments using comparative analysis via automatic measures and statistics.
The evaluation panel provides graphs, tests for statistical significance and n-gram statistics.
We also present a demo server http://wmt.ufal.cz with WMT14 and WMT15 translations
Results of the WMT15 Metrics Shared Task
This paper presents the results of the WMT15 Metrics Shared Task. We asked
participants of this task to score the outputs of the MT systems involved in
the WMT15 Shared Translation Task. We collected scores of 46 metrics from 11
research groups. In addition to that, we computed scores of 7 standard metrics
(BLEU, SentBLEU, NIST, WER, PER, TER and CDER) as baselines. The collected scores were
evaluated in terms of system level correlation (how well each metric's scores
correlate with WMT15 official manual ranking of systems) and in terms of segment
level correlation (how often a metric agrees with humans in comparing two
translations of a particular sentence)