Charles University

Biblio at Institute of Formal and Applied Linguistics
Not a member yet
    539 research outputs found

    Building a Bilingual ValLex Using Treebank Token Alignment: First Observations

    No full text
    In this paper we explore the potential and limitations of a concept of building a bilingual valency lexicon based on the alignment of nodes in a parallel treebank. Our aim is to build an electronic CzechEnglish Valency Lexicon by collecting equivalences from bilingual treebank data and storing them in two already existing electronic valency lexicons, PDT-VALLEX and Engvallex. For this task a special annotation interface has been built upon the TrEd editor, allowing quick and easy collecting of frame equivalences in either of the source lexicons. The issues questioning the annotation practice encountered during the first months of annotation include limitations of technical character, theory-dependent limitations and limitations concerning the achievable degree of quality of human annotation. The issues of special interest for both linguists and MT specialists involved in the project include linguistically motivated non-balance between the frame equivalents, either in number or in type of valency participants. The first phases of annotation so far attest the assumption that there is a unique correspondence between the functors of the translation-equivalent frames. Also, hardly any linguistically significant non-balance between the frames has been found, which is partly promising considering the linguistic theory used and partly caused by little stylistic variety of the annotated corpus texts

    MT via Deep Syntax

    Get PDF

    Tackling Sparse Data Issue in Machine Translation Evaluation

    No full text
    We illustrate and explain problems of n-grams-based machine translation (MT) metrics (e.g. BLEU) when applied to morphologically rich languages such as Czech. A novel metric SemPOS based on the deep-syntactic representation of the sentence tackles the issue and retains the performance for translation to English as well

    Data Issues in English-to-Hindi Machine Translation

    No full text
    Statistical machine translation to morphologically richer languages is a challenging task and more so if the source and target languages differ in word order. Current state-of-the-art MT systems thus deliver mediocre results. Adding more parallel data often helps improve the results; if it doesn't, it may be caused by various problems such as different domains, bad alignment or noise in the new data. In this paper we evaluate the English-to-Hindi MT task from this data perspective. We discuss several available parallel data sources and provide cross-evaluation results on their combinations using two freely available statistical MT systems. Together with the error analysis, we also present a new tool for viewing aligned corpora, which makes it easier to detect difficult parts in the data even for a developer not speaking the target language

    2010 Failures in English-Czech Phrase-Based MT

    No full text
    The paper describes our experiments with English-Czech machine translation for WMT10 in 2010. Focusing primarily on the translation to Czech, our additions to the standard Moses phrase-based MT pipeline include two-step translation to overcome target-side data sparseness and optimization towards SemPOS, a metric better suited for evaluating Czech. Unfortunately, none of the approaches bring a significant improvement over our standard setup

    Automatic Source Code Reduction

    No full text
    The aim of this paper is to introduce Reductor, a program that automatically removes unused parts of the source code of valid programs written in the Mercury language. Reductor implements two main kinds of reductions: statical reduction and dynamical reduction. In the statical reduction, Reductor exploits semantic analysis of the Melbourne Mercury Compiler to nd routines which can be removed from the program. Dynamical reduction of routines additionally uses Mercury Deep Profiler and some sample input data for the program to remove unused contents of the program routines. Reductor modifies the sources of the program in a way, which keeps the formatting of the original program source so that the reduced code is further editable

    Jak se dělá strojový překlad

    No full text
    An introduction to machine translation for the participants of seminar led by Lucie Mladová

    The Lexical Population of Semantic Types in Hanks’s PDEV

    No full text
    This contribution reports on an ongoing analysis of the Pattern Dictionary of English Verbs (PDEV, Hanks 2007b) with respect to both its consistency and reproducibility of use by different users. We address, in particular, the assign-ment of Semantic Type labels to noun collocates of verbs, in a series of experi-ments conducted at the Institute of Formal and Applied Linguistics of the Charles University in Prague

    Ways of Evaluation of the Annotators in Building the Prague Czech-English Dependency Treebank

    No full text
    The paper presents several ways to measure and evaluate the annotation and annotators, proposed and used during the building of the Czech part of the Prague Czech-English Dependency Treebank. At first, the basic principles of the treebank annotation project are introduced (division to three layers: morphological, analytical and tectogrammatical). The main part of the paper describes in detail one of the important phases of the annotation process: three ways of evaluation of the annotators - inter-annotator agreement, error rate and performance. The measuring of the inter-annotator agreement is complicated by the fact that the data contain added and deleted nodes, making the alignment between annotations non-trivial. The error rate is measured by a set of automatic checking procedures that guard the validity of some invariants in the data. The performance of the annotators is measured by a booking web application. All three measures are later compared and related to each other

    Aim and result – A Swedish-Czech comparison of consecutive clauses

    No full text
    Three types of subordinate clauses express the dependence of the event they denote on an event that is denoted by the governing clause: these are the result clause (henceforth RC), the purpose clause (henceforth PC), and the minimum requirement clause (henceforth MRC). This paper analyzes the semantics of various forms of event consecutiveness expressed by these three clause types in Swedish (mainly based on Clausén et al., 2003, pp. 633-639) and Czech (besed on Karlík et al., 1995), respectively. Findings in a 2-million Swedish-Czech parallel corpus suggest that different languages may exhibit different clause-type preferences (RC/PC/MRC) when referring to consecutive events of the same type

    58

    full texts

    539

    metadata records
    Updated in last 30 days.
    Biblio at Institute of Formal and Applied Linguistics
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇