539 research outputs found
Sort by
Building a Bilingual ValLex Using Treebank Token Alignment: First Observations
In this paper we explore the potential and limitations of a concept of building a bilingual valency lexicon based on the alignment of nodes
in a parallel treebank. Our aim is to build an electronic CzechEnglish Valency Lexicon by collecting equivalences from bilingual
treebank data and storing them in two already existing electronic valency lexicons, PDT-VALLEX and Engvallex. For this task a special
annotation interface has been built upon the TrEd editor, allowing quick and easy collecting of frame equivalences in either of the source
lexicons. The issues questioning the annotation practice encountered during the first months of annotation include limitations of technical
character, theory-dependent limitations and limitations concerning the achievable degree of quality of human annotation. The issues of
special interest for both linguists and MT specialists involved in the project include linguistically motivated non-balance between the
frame equivalents, either in number or in type of valency participants. The first phases of annotation so far attest the assumption that
there is a unique correspondence between the functors of the translation-equivalent frames. Also, hardly any linguistically significant
non-balance between the frames has been found, which is partly promising considering the linguistic theory used and partly caused by
little stylistic variety of the annotated corpus texts
Tackling Sparse Data Issue in Machine Translation Evaluation
We illustrate and explain problems of
n-grams-based machine translation (MT)
metrics (e.g. BLEU) when applied to
morphologically rich languages such as
Czech. A novel metric SemPOS based
on the deep-syntactic representation of the
sentence tackles the issue and retains the
performance for translation to English as
well
Data Issues in English-to-Hindi Machine Translation
Statistical machine translation to morphologically richer languages is a challenging task and more so if the source and target languages differ in word order. Current
state-of-the-art MT systems thus deliver mediocre results. Adding more parallel data often helps improve the results; if it doesn't, it may be caused by various problems such as different domains, bad alignment or noise in the new data.
In this paper we evaluate the English-to-Hindi MT task from this data perspective. We discuss several available parallel data sources and provide cross-evaluation results on their combinations using two freely available
statistical MT systems. Together with the error analysis, we also present a new tool
for viewing aligned corpora, which makes it easier to detect difficult parts in the data even for a developer not speaking the target
language
2010 Failures in English-Czech Phrase-Based MT
The paper describes our experiments with
English-Czech machine translation for
WMT10 in 2010. Focusing primarily
on the translation to Czech, our additions
to the standard Moses phrase-based MT
pipeline include two-step translation to
overcome target-side data sparseness and
optimization towards SemPOS, a metric
better suited for evaluating Czech. Unfortunately,
none of the approaches bring a
significant improvement over our standard
setup
Automatic Source Code Reduction
The aim of this paper is to introduce Reductor,
a program that automatically removes unused parts of the source code of valid programs written in the Mercury language. Reductor implements two main kinds of reductions: statical reduction and dynamical reduction. In the statical reduction, Reductor exploits semantic analysis of the Melbourne Mercury Compiler to nd routines which can be removed from the program. Dynamical reduction of routines additionally uses Mercury Deep Profiler and some sample input data for the program to remove unused contents of the program routines. Reductor modifies the sources of the program in a way, which keeps the formatting of the original program source so that the reduced code is further editable
Jak se dělá strojový překlad
An introduction to machine translation for the participants of seminar led by Lucie Mladová
The Lexical Population of Semantic Types in Hanks’s PDEV
This contribution reports on an ongoing analysis of the Pattern Dictionary of English Verbs (PDEV, Hanks 2007b) with respect to both its consistency and reproducibility of use by different users. We address, in particular, the assign-ment of Semantic Type labels to noun collocates of verbs, in a series of experi-ments conducted at the Institute of Formal and Applied Linguistics of the Charles University in Prague
Ways of Evaluation of the Annotators in Building the Prague Czech-English Dependency Treebank
The paper presents several ways to measure and evaluate the annotation and annotators, proposed and used during the building of the Czech part of the Prague Czech-English Dependency Treebank. At first, the basic principles of the treebank annotation project are introduced (division to three layers: morphological, analytical and tectogrammatical). The main part of the paper describes in detail one of the important phases of the annotation process: three ways of evaluation of the annotators - inter-annotator agreement, error rate and performance. The measuring of the inter-annotator agreement is complicated by the fact that the data contain added and deleted nodes, making the alignment between annotations non-trivial. The error rate is measured by a set of automatic checking procedures that guard the validity of some invariants in the data. The performance of the annotators is measured by a booking web application. All three measures are later compared and related to each other
Aim and result – A Swedish-Czech comparison of consecutive clauses
Three types of subordinate clauses express the dependence of the event they denote on an event that is denoted by the governing clause: these are the result clause (henceforth RC), the purpose clause (henceforth PC), and the minimum requirement clause (henceforth MRC). This paper analyzes the semantics of various forms of event consecutiveness expressed by these three clause types in Swedish (mainly based on Clausén et al., 2003, pp. 633-639) and Czech (besed on Karlík et al., 1995), respectively. Findings in a 2-million Swedish-Czech parallel corpus suggest that different languages may exhibit different clause-type preferences (RC/PC/MRC) when referring to consecutive events of the same type