539 research outputs found
Sort by
Evaluating Quality of Machine Translation from Czech to Slovak
We focus on machine translation between closely related
languages, in particular Czech and Slovak. We mention the specifics of
evaluating MT quality for closely related languages. The main
contribution is the test and confirmation of the old assumption that
rule-based systems with shallow transfer still work better than
current state-of-the-art statistical systems in this setting, unless
huge amounts of data are availabl
Word Alignment as a Combinatorial Task
An introduction to the task of word alignment for colleagues at the Department of applied mathematics (KAM)
Improving Translation Model by Monolingual Data
We use target-side monolingual data to extend
the vocabulary of the translation model
in statistical machine translation. This method
called “reverse self-training” improves the decoder’s
ability to produce grammatically correct
translations into languages with morphology
richer than the source language esp. in
small-data setting. We empirically evaluate
the gains for several pairs of European
languages and discuss some approaches of
the underlying back-off techniques needed to
translate unseen forms of known words. We
also provide a description of the systems we
submitted to WMT11 Shared Task
Analyzing Error Types in English-Czech Machine Translation
This paper examines two techniques of manual evaluation that can be used to identify error
types of individual machine translation systems. The first technique of “blind post-editing” is
being used in WMT evaluation campaigns since 2009 and manually constructed data of this
type are available for various language pairs. The second technique of explicit marking of errors
has been used in the past as well.
We propose a method for interpreting blind post-editing data at a finer level and compare
the results with explicit marking of errors. While the human annotation of either of the techniques
is not exactly reproducible (relatively low agreement), both techniques lead to similar
observations of differences of the systems. Specifically, we are able to suggest which errors in
MT output are easy and hard to correct with no access to the source, a situation experienced by
users who do not understand the source language
Quiz-Based Evaluation of Machine Translation
This paper proposes a new method of manual evaluation for statistical machine translation,
the so-called quiz-based evaluation, estimating whether people are able to extract information
from machine-translated texts reliably. We apply the method to two commercial and two experimental
MT systems that participated in WMT 2010 in English-to-Czech translation. We
report inter-annotator agreement for the evaluation as well as the outcomes of the individual
systems. The quiz-based evaluation suggests rather different ranking of the systems compared
to the WMT 2010 manual and automatic metrics. We also see that overall, MT quality is becoming
acceptable for obtaining information from the text: about 80% of questions can be answered
correctly given only machine-translated text
Combining Diverse Word-Alignment Symmetrizations Improves Dependency Tree Projection
For many languages, we are not able to train any supervised parser, because there are
no manually annotated data available. This problem can be solved by using a parallel corpus
with English, parsing the English side, projecting the dependencies through word-alignment
connections, and training a parser on the projected trees. In this paper, we introduce a
simple algorithm using a combination of various word-alignment symmetrizations. We prove
that our method outperforms previous work, even though it uses McDonald's maximum-spanning-tree
parser as it is, without any "unsupervised" modifications
Prague Czech-English Dependency Treebank 2.0
The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part
Addicter 2.0
Addicter stands for Automatic Detection and DIsplay of Common Translation ERrors. It is a set of tools (mostly scripts written in Perl) that help with error analysis for machine translation. The second version contains a greatly improved viewer and many new modules for automatic detection and classification of errors
Dependency Parsing
Dependency parsing has been a prime focus of NLP research of late due to its
ability to help parse languages with a free word order. Dependency parsing has been shown
to improve NLP systems in certain languages and in many cases is considered the state of
the art in the field. The use of dependency parsing has mostly been limited to free word
order languages, however the usefulness of dependency structures may yield improvements
in many of the word’s 6,000+ languages.
I will give an overview of the field of dependency parsing while giving my aims for
future research. Many NLP applications rely heavily on the quality of dependency parsing.
For this reason, I will examine how different parsers and annotation schemes influence the
overall NLP pipeline in regards to machine translation as well as the the baseline parsing
accuracy
Tamil Dependency Treebank (TamilTB) - 0.1 Annotation Manual
Tamil Dependency Treebank (TamilTB) is
an attempt to develop a syntactically annotated corpora for Tamil. TamilTB
contains 600 sentences enriched with manual annotation of morphology and
dependency syntax in the style of Prague Dependency Treebank. TamilTB
has been created at the Institute of Formal and Applied Linguistics, Charles
University in Prague. This report serves the purpose of how the annotation has been done at morphological level and syntactic level. Annotation scheme has been elaborately discussed with examples