539 research outputs found
Sort by
Česko-slovenský paralelný korpus
Czech-Slovak parallel corpus consisting of several freely available corpora. Corpus is given in both plaintext format and with an automatic morphological annotation
English-Slovak Parallel Corpus
English-Slovak parallel corpus consisting of several freely available corpora. Corpus is given in both in plaintext format and with an automatic morphological annotatio
A database of semantic clusters of verb usages
We are presenting VPS-30-En, a small lexical resource that contains the following 30 English verbs: access, ally, arrive, breathe,
claim, cool, crush, cry, deny, enlarge, enlist, forge, furnish, hail, halt, part, plough, plug, pour, say, smash, smell, steer, submit, swell,
tell, throw, trouble, wake and yield. We have created and have been using VPS-30-En to explore the interannotator agreement potential
of the Corpus Pattern Analysis. VPS-30-En is a small snapshot of the Pattern Dictionary of English Verbs (Hanks and Pustejovsky,
2005), which we revised (both the entries and the annotated concordances) and enhanced with additional annotations. It is freely
available at http://ufal.mff.cuni.cz/spr. In this paper, we compare the annotation scheme of VPS-30-En with the original PDEV. We
also describe the adjustments we have made and their motivation, as well as the most pervasive causes of interannotator disagreements
The Study of Effect of Length in Morphological Segmentation of Agglutinative Languages
Morph length is one of the indicative feature
that helps learning the morphology of languages,
in particular agglutinative languages.
In this paper, we introduce a simple unsupervised
model for morphological segmentation
and study how the knowledge of morph
length affect the performance of the segmentation
task under the Bayesian framework.
The model is based on (Goldwater et
al., 2006) unigram word segmentation model
and assumes a simple prior distribution over
morph length. We experiment this model
on two highly related and agglutinative languages
namely Tamil and Telugu, and compare
our results with the state of the art Morfessor
system. We show that, knowledge of
morph length has a positive impact and provides
competitive results in terms of overall
performance
Maintaining consistency of monolingual verb entries with interannotator agreement
There is no objectively correct way to create a monolingual entry of a polysemous verb. By structuring a verb into readings, we impose our conception onto lexicon users, no matter how big a corpus we use in support. How do we make sure that our structuring is intelligible for others?
We are performing an experiment with the validation of the fully corpus-based Pattern Dictionary of English Verbs (Hanks & Pustejovsky, 2005), created according to the lexical theory Corpus Pattern Analysis (CPA). The lexicon is interlinked with a large corpus, in which several hundred randomly selected concordances of each processed verb are manually annotated with numbers of their corresponding lexicon readings (“patterns”). It would be interesting to prove (or falsify) the leading assumption of CPA that, given the patterns are based on a large corpus, individual introspection has been minimized and most people can agree on this particular semantic structuring. We have encoded the guidelines for assigning concordances to patterns and hired annotators to annotate random samples of verbs cotained in the lexicon. Apart from measuring the interannotator agreement, we analyze and adjudicate the disagreements. The outcome is offered to the lexicographer as feedback. The lexicographer revises his entries and the agreement can be measured againg on a different random sample to test whether or not the revision has brought an improvement of the interannotator agreement score. A high interannotator agreement suggests that lexicon users are likely to find a pattern corresponding to a random verb use of which they seek explanation. A low agreement score gives a warning that there are patterns missing or vague.
We focus on machine-learning applications, but we believe that this procedure is of interest even for quality management in human lexicography
Unsupervised Dependency Parsing using Reducibility and Fertility features
Popis systemu nerizeneho zavislostniho parsingu zalozeneho na Gibbsove samplingu. Novy pristup predstavuje vlastnosti vypusttelnsti a fertility
Prague Dependency Style Treebank for Tamil
Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the
language itself. Only very few languages possess these scarce resources. In this paper, we describe our efforts in syntactically annotating
a small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT) and consists of
annotation at 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce
annotation schemes i.e. positional tagging for m-layer and dependency relations for a-layers. Finally, we discuss some of the issues in
treebank development for Tamil
Ensemble Parsing and its Effect on Machine Translation
The focus of much of dependency parsing is on creating new modeling techniques and examining new feature sets for existing dependency models. Often these new models are lucky to achieve equivalent results with the current state of
the art results and often perform worse. These approaches are for languages that are often resource-rich and have ample training data available for dependency parsing. For this reason, the accuracy scores are often quite high. This, by its very nature, makes it quite difficult to create a significantly large increase in the current state-of-the-art. Research in this area is often concerned with small accuracy changes or very specific localized changes, such as increasing accuracy of a particular linguistic construction. With so many modeling techniques available to languages with large resources the problem exists on how to exploit the current techniques with the use of combination, or ensemble, techniques along with this plethora of data.
Dependency parsers are almost ubiquitously evaluated on their accuracy scores, these scores say nothing of the complexity and usefulness of the resulting structures. The structures may have more complexity due to the depth of their co-
ordination or noun phrases. As dependency parses are basic structures in which other systems are built upon, it would seem more reasonable to judge these parsers down the NLP pipeline. The types of parsing errors that cause significant
problems in other NLP applications is currently an unknown
Automatic MT Error Analysis: Hjerson Helping Addicter
We present a complex, open source tool for detailed machine translation error analysis providing the user with automatic error detection
and classification, several monolingual alignment algorithms as well as with training and test corpus browsing. The tool is the result of
a merge of automatic error detection and classification of Hjerson (Popović, 2011) and Addicter (Zeman et al., 2011) into the pipeline
and web visualization of Addicter. It classifies errors into categories similar to those of Vilar et al. (2006), such as: morphological, reordering, missing words, extra words and lexical errors. The graphical user interface shows alignments in both training corpus and
test data; the different classes of errors are colored. Also, the summary of errors can be displayed to provide an overall view of the MT
system’s weaknesses. The tool was developed in Linux, but it was tested on Windows too
Khresmoi: Multimodal Multilingual Medical Information Search
Khresmoi is a European Integrated Project developing a multilingual multimodal search and access system for medical and health information and documents. It addresses the challenges of searching through huge amounts of medical data, including general medical information available on the internet, as well as radiology data in hospital archives. It is developing novel semantic search and visual search techniques for the medical domain. At the MIE Village of the
Future, Khresmoi proposes to have two interactive demonstrations of the system under development, as well as an overview oral presentation and potentially some poster presentations