539 research outputs found
Sort by
Srovnání koreferenčních výrazů v češtině a angličtině
In this work, we present a comprehensive study on correspondences between certain classes of coreferential expressions in English and Czech. We focus on central pronouns, relative pronouns, and anaphoric zeros. We designed an alignment-refining algorithm for English personal and possessive pronouns and Czech relative pronouns that improves the quality of alignment links not only for the classes it aimed at but also in general. Moreover, the instances of anaphoric expressions we focus on were manually annotated with their alignment counterparts, which served as a basis for this empirical study. The collected statistics of correspondences are contrasted with theoretical assumptions regarding the use of anaphoric means in the languages under analysis, such as pro-drop properties, the use of finite and non-finite constructions, etc. Finally, we present the ways how the observed correspondences can be exploited in cross-lingual coreference resolution
KLcpos3 - a Language Similarity Measure for Delexicalized Parser Transfer
We present KLcpos3, a language similarity measure based on Kullback-Leibler
divergence of coarse part-of-speech tag trigram distributions in tagged
corpora. It has been designed for multilingual delexicalized parsing, both for
source treebank selection in single-source parser transfer, and for source
treebank weighting in multi-source transfer. In the selection task, KLcpos3
identifies the best source treebank in 8 out of 18 cases. In the weighting
task, it brings +4.5% UAS absolute, compared to unweighted parse tree
combination
MSTperl parser (2015-05-19)
MSTperl is a Perl reimplementation of the MST parser of Ryan McDonald, with several additional advanced functions, such as support for parallel features
Comparison of Coreference Resolvers for Deep Syntax Translation
This work focuses on using anaphora for machine translation with deep-syntactic transfer. We compare multiple coreference resolvers for English in terms of how they affect the quality of pronoun translation in English-Czech and English-Dutch machine translation systems with deep transfer. We examine which pronouns in the target language depend on anaphoric information, and design rules that take advantage of this information. The resolvers’ performance measured by translation quality is contrasted with their
intrinsic evaluation results. In addition, a more detailed manual analysis of English-to-Czech translation was carried out
Coreference chains in Czech, English and Russian: Preliminary findings
Tento článek je pilotní srovnavací výzkum koreferenčních řetězců v češtině, angličtině a ruštině. Podrobili jsme analýze 16 srovnatelných textů ve třech jazycích. Naší motivací bylo zjistit lingvistickou strukturu koreferenčních řetězců v těchto jazycích a určit, které faktory ovlivňují tuto strukturu
The LINDAT/CLARIN large infrastructure: data and services
LINDAT/CLARIN, as a node of the pan-European research infrastructure Clarin ERIC, has been presented. Its repository has been featured together with data archivation techniques and related web services and applications
WMT2013 Test Set in Vietnamese
The file contains the standard WMT 2013 test set, manually translated from English to Vietnamese, thus extending the WMT 2013 multi-parallel set of languages (English, Czech, German, French, Spanish, Russian)
Deep Linguistic Information in Machine Translation
Fundamentals of Machine Translation technology using Deep language analysis using the traditional "Vauquois triangle" have been presented, which now use advanced statistical and machine learning techniques. In addition, latest WMT 2015 Shared Task competition results for the en-cs pair have been shown in which the Chimera system created at UFAL MFF UK won
Retrieving information about medical symptoms
This paper details methods, results and analysis of the CLEF 2015 eHealth Evaluation Lab, Task 2. This task investigates the effectiveness of web search engines in providing access to medical information with the aim of fostering advances in the development of these technologie
Abstract Meaning Representation: comparing languages and their representations
Abstract Meaning Representation is a newly developed formalism for representing meaning, which abstracts from syntax and some other phenomena but is (still) language-dependent. It has been developed by a consortium of mostly U.S. universities, with team members including Martha Palmer, Kevin Knight, Philipp Koehn, Ulf Hermjakob, Kathy McKeown, Nianwen Xue and others. In the talk, the basic facts about the AMR will be presented, and then comparison will be made for Czech and English as carried out in detail on a small 100-sentence corpus; some examples from Chinese-English comparison will also be shown. In addition, AMR will be compared to the deep syntactic representation used in the set of Prague Dependency Treebanks, and observations will be made about the level of abstraction used in these two formalisms. There is an ongoing work on possible conversion between the Prague deep syntactic representation and the AMR representation, and the main issues of such a conversion will also be described, as well as plans for future studies and possible corpus annotation work in the nearest future