Gonzaga University Repository
Not a member yet
2718 research outputs found
Sort by
Word, Lemma, and Part of Speech (Cuba)
Corpus del Español word, lemma, and part of speech data format.
Zip folder contains 20 .txt files of linguistic data from Cuba split into two categories: General (g) and Blogs (b).
Texts are separated by a line with ## and the textID
Word, Lemma, and Part of Speech (El Salvador)
Corpus del Español word, lemma, and part of speech data format.
Zip folder contains 20 .txt files of linguistic data from El Salvador split into two categories: General (g) and Blogs (b).
Texts are separated by a line with ## and the textID
Word, Lemma, and Part of Speech (Paraguay)
Corpus del Español word, lemma, and part of speech data format.
Zip folder contains 20 .txt files of linguistic data from Paraguay split into two categories: General (g) and Blogs (b).
Texts are separated by a line with ## and the textID
Word, Lemma, and Part of Speech (Uruguay)
Corpus del Español word, lemma, and part of speech data format.
Zip folder contains 20 .txt files of linguistic data from Uruguay split into two categories: General (g) and Blogs (b).
Texts are separated by a line with ## and the textID
Linear Text (Venezuela)
Corpus del Español linear text data format for texts from Venezuela. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data
Linear Text (El Salvador)
Corpus del Español linear text data format for texts from El Salvador. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data
Linear Text (Panamá)
Corpus del Español linear text data format for texts from Panamá. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data
Database (Venezuela)
Corpus del Español database format. This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL. Zip folder contains 20 .txt files of linguistic data from Venezuela split into two categories: General (g) and Blogs (b). See Full-text corpus data for more information on how to use the database format
Corpus del Español Complete Database
Complete Corpus del Español database format for linguistic data from 21 Spanish speaking countries. This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL.
This TAR file includes 21 zipped folders, each containing 20 .txt files of linguistic data split into two categories: General (g) and Blogs (b).
See Full-text corpus data for more information on how to use the database format
NOW 2016-03 March
Corpus of News on the Web data for March 2016.
The TAR folder contains linguistic data in three formats: Database: This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL. See Full-text corpus data for more information on how to use the database format. Linear Text: This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Word, Lemma, Part of Speech: Texts are separated by a line with ## and the textID