539 research outputs found
Sort by
Language Richness of the Web
We have built a corpus containing texts in 106 languages from texts available on the Internet and on Wikipedia. The W2C Web Corpus contains 54.7 GB of text and the W2C Wiki Corpus contains 8.5 GB of text. The W2C Web Corpus contains more than 100 MB of text available for 75 languages. At least 10 MB of text is available for 100 languages. These corpora are a unique data source for linguists, since they outclass all published works both in the size of the material collected and the number of languages covered. This language data resource can be of use particularly to researchers specialized in multilingual technologies development. We also developed software that greatly simplifies the creation of a new text corpus for a given language, using text materials freely available on the Internet. Special attention was given to components for filtering and de-duplication that allow to keep the material quality very high
Managing Uncertainty in Semantic Tagging
Low interannotator agreement (IAA) is a
well-known issue in manual semantic tagging
(sense tagging). IAA correlates with
the granularity of word senses and they
both correlate with the amount of information
they give as well as with its reliability.
We compare different approaches to semantic
tagging in WordNet, FrameNet, Prop-
Bank and OntoNotes with a small tagged
data sample based on the Corpus Pattern
Analysis to present the reliable information
gain (RG), a measure used to optimize the
semantic granularity of a sense inventory
with respect to its reliability indicated by
the IAA in the given data set. RG can also
be used as feedback for lexicographers, and
as a supporting component of automatic semantic
classifiers, especially when dealing
with a very fine-grained set of semantic categories
Simple and Effective Parameter Tuning for Domain Adaptation of Statistical Machine Translation
Current state-of-the-art Statistical Machine Translation systems are based on log-linear models
that combine a set of feature functions to score translation hypotheses during decoding. The
models are parametrized by a vector of weights usually optimized on a set of sentences and
their reference translations, called development data. In this paper, we explore a (common
and industry relevant) scenario where a system trained and tuned on general domain data
needs to be adapted to a specific domain for which no or only very limited in-domain bilingual
data is available. It turns out that such systems can be adapted successfully by re-tuning model
parameters using surprisingly small amounts of parallel in-domain data, by cross-tuning or no
tuning at all. We show in detail how and why this is effective, compare the approaches and
effort involved. We also study the effect of system hyperparameters (such as maximum phrase
length and development data size) and their optimal values in this scenario
Hybrid Combination of Constituency and Dependency Trees into an Ensemble Dependency Parser
Dependency parsing has made many advancements in recent years, in particular for English. There are a few dependency parsers that achieve comparable accuracy scores with each other but with very different types of errors. This paper examines creating a new dependency structure through ensemble learning using a hybrid of the outputs of various parsers. We combine all tree outputs into a weighted edge graph, using 4 weighting mechanisms. The weighted edge graph is the input into our ensemble system and is a hybrid of very different parsing techniques (constituent parsers, transition-based dependency parsers, and a graph-based parser). From this graph we take a maximum spanning tree. We examine the new dependency structure in terms of accuracy and errors on individual part-of-speech values.
The results indicate that using a greater number of more varied parsers will improve accuracy results. The combined ensemble system, using 5 parsers based on 3 different parsing techniques, achieves an accuracy score of 92.58%, beating all single parsers on the Wall Street Journal section 23 test set. Additionally, the ensemble system reduces the average relative error on selected POS tags by 9.82%
Building parallel corpora through social network gaming
Building training data is labor-intensive and presents a major obstacle to the advancement of Natural Language Processing (NLP)
systems. A prime use of NLP technologies has been toward the construction machine translation systems. The most common form of
machine translation systems are phrase based systems that require extensive training data. Building this training data is both expensive
and error prone. Emerging technologies, such as social networks and serious games, offer a unique opportunity to change how we
construct training data. These serious games, or games with a purpose, have been constructed for sentence segmentation, image labeling,
and co-reference resolution. These games work on three levels: They provide entertainment to the players, the reinforce information
the player might be learning, and they provide data to researchers. Most of these systems while well intended and well developed, have
lacked participation.
We present, a set of linguistically based games that aim to construct parallel corpora for a multitude of languages and allow players to
start learning and improving their own vocabulary in these languages. As of the first release of the games, GlobeOtter is available on
Facebook as a social network game. The release of this game is meant to change the default position in the field, from creating games
that only linguists play, to releasing linguistic games on a platform that has a natural user base and ability to grow
The Joy of Parallelism with CzEng 1.0
CzEng 1.0 is an updated release of our Czech-English parallel corpus, freely
available for non-commercial research or educational purposes. In this
release, we approximately doubled
the corpus size, reaching 15 million sentence
pairs (about 200 million tokens per language). More importantly, we carefully
filtered the data to reduce the amount of non-matching sentence pairs.
CzEng 1.0 is automatically aligned at the level of sentences as well as words.
We provide not only the plain text representation, but also automatic
morphological tags, surface syntactic as well as deep syntactic dependency parse
trees and automatic co-reference links in both English and Czech.
This paper describes key properties of the released resource including the
distribution of text domains,
the corpus data formats, and a toolkit to handle the provided rich annotation. We also
summarize the procedure of the rich annotation (incl. co-reference
resolution) and of the automatic filtering. Finally, we provide some suggestions
on exploiting such an automatically annotated sentence-parallel corpus
WMT 2011 Testing Set in Slovak
WMT 2011 Workshop testing set manually translated from Czech and English into Slovak
Terra: a Collection of Translation Error-Annotated Corpora
Recently the first methods of automatic diagnostics of machine translation have emerged; since this area of research is relatively young,
the efforts are not coordinated. We present a collection of translation error-annotated corpora, consisting of automatically produced trans-
lations and their detailed manual translation error analysis. Using the collected corpora we evaluate the available state-of-the-art methods
of MT diagnostics and assess, how well the methods perform, how they compare to each other and whether they can be useful in practice
Manually ranked outputs of Czech-Slovak translations
Data from three sources (part of Acquis, WMT test set and sentences selected from the set of books) translated by 5 machine translation systems (Česílko, Česílko 2, Google Translate and Moses with different settings) from Czech to Slovak and evaluated by three annotators. The translations were manually ordered according to their quality