CLARIN-PL
Not a member yet
504 research outputs found
Sort by
Vector Extractor
Collocations presented are based on co-occurrences of a selected noun with several features describing it and linked with it by syntactic dependencies. The recognised features are: modification by an adjective (AdjMod), modification by a noun in genitive (NGenMod), coordination with a noun (NCoord) and linking to a verb as its subject (VSubj)
HaskEN
HaskEN is an English phraseological database designed for language professionals including linguists, language teachers, lexicographers, language materials developers and translators. Query results can be visualised and exported as spreadsheets
TaKIPI
TaKIPI is a tagger of Polish language that is a tool which assigns morpho-syntactic markers to words in the text.
The tagger assumes a morpho-syntactic description of IPI PAN corpus tagset. Contextual disambiguation is carried out via a small set of hand-written rules and via a bigger number of rules automatically extracted by means of the algorithm of the induction of decision trees C4.5. During the process of tagger's learning and functioning, the context of each word's occurence in the text is represented as a feature vector of a constant length. Such vector is obtained by means of hand-written functional expressions of JOSKIPI formalism, which refer to morpho-syntactic properties of the context
plWordNet as the Cornerstone of a Toolkit of Lexico-semantic Resources
A wordnet is many things to many people: a graph of inter-related lexicalised concepts, a taxonomy, a thesaurus, and so on. A wordnet makes good sense as the mainstay of any deep automated semantic analysis of text. We have begun the construction of a multi-component, multi-use toolkit of natural language processing tools with plWordNet, a very large Polish wordnet, at its centre. The components will include plWordNet and its mapping onto an ontology (the upper level and elements of the middle level), a lexicon of proper names and a semantic valency lexicon. Some of those elements will be aligned with plWordNet, and there will be a mapping onto Princeton WordNet. Several challenging applications will show the utility of the toolkit in practice
The system of register labels in plWordNet v. 5 (Guidelines)
The pdf document contains guidelines of the description of the register of lexical units in the polish part of plWordNe
Lists of semantic relatedness
Dystrybucyjne Podobieństwo Semantyczne (DPS, ang. Measure of Semantic Relatedness) obrazuje podobieństwo pomiędzy parami wyrazów na podstawie analizy ich współwystępowania w korpusach tekstów. Ogólną sposób wydobywania podobieństwa można przedstawić następująco. W pierwszej kolejności wszystkie konkteksty interesujących słów są analizowane pod kątem współwystępowania z innymi słowami
Spokes search engine for Polish conversational data
Spokes is an online service for conversational corpus data search and exploration as part of the Polish CLARIN infrastructure. The underlying corpus contains more than 2 million words of time-aligned transcriptions of casual spoken discourse. The service is available both as a web application and as a REST service
Liner2.4
A framework for multitask sequence labeling dedicated for natural language processing tasks