Gonzaga University Repository
Not a member yet
    2718 research outputs found

    Word, Lemma, and Part of Speech (Cuba)

    No full text
    Corpus del Español word, lemma, and part of speech data format. Zip folder contains 20 .txt files of linguistic data from Cuba split into two categories: General (g) and Blogs (b). Texts are separated by a line with ## and the textID

    Word, Lemma, and Part of Speech (El Salvador)

    No full text
    Corpus del Español word, lemma, and part of speech data format. Zip folder contains 20 .txt files of linguistic data from El Salvador split into two categories: General (g) and Blogs (b). Texts are separated by a line with ## and the textID

    Word, Lemma, and Part of Speech (Paraguay)

    No full text
    Corpus del Español word, lemma, and part of speech data format. Zip folder contains 20 .txt files of linguistic data from Paraguay split into two categories: General (g) and Blogs (b). Texts are separated by a line with ## and the textID

    Word, Lemma, and Part of Speech (Uruguay)

    No full text
    Corpus del Español word, lemma, and part of speech data format. Zip folder contains 20 .txt files of linguistic data from Uruguay split into two categories: General (g) and Blogs (b). Texts are separated by a line with ## and the textID

    Linear Text (Venezuela)

    No full text
    Corpus del Español linear text data format for texts from Venezuela. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data

    Linear Text (El Salvador)

    No full text
    Corpus del Español linear text data format for texts from El Salvador. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data

    Linear Text (Panamá)

    No full text
    Corpus del Español linear text data format for texts from Panamá. This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Zipped folder contains 20 .txt files of General (g) and Blog (b) data

    Database (Venezuela)

    No full text
    Corpus del Español database format. This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL. Zip folder contains 20 .txt files of linguistic data from Venezuela split into two categories: General (g) and Blogs (b). See Full-text corpus data for more information on how to use the database format

    Corpus del Español Complete Database

    No full text
    Complete Corpus del Español database format for linguistic data from 21 Spanish speaking countries. This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL. This TAR file includes 21 zipped folders, each containing 20 .txt files of linguistic data split into two categories: General (g) and Blogs (b). See Full-text corpus data for more information on how to use the database format

    NOW 2016-03 March

    No full text
    Corpus of News on the Web data for March 2016. The TAR folder contains linguistic data in three formats: Database: This is the format allows for the most robust searches and allows for powerful JOINs across corpus, lexicon, and source tables but requires knowledge of SQL. See Full-text corpus data for more information on how to use the database format. Linear Text: This format provides a textID for each text, and then the entire text on the same line. In this format, words are not annotated for part of speech or lemma. In addition, contracted words like can\u27t are separated into two parts (ca n\u27t) and punctuation is separated from words (eye level . As her). Word, Lemma, Part of Speech: Texts are separated by a line with ## and the textID

    1,331

    full texts

    2,718

    metadata records
    Updated in last 30 days.
    Gonzaga University Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇