1,721,203 research outputs found

    Eye dialect: translating the untranslatable

    Get PDF
    The term ‘eye dialect’ was first coined in 1925 by George P. Krapp inThe English Language in America(McArthur 1998). The term was used to describe the phenomenon of unconventional spelling used to reproduce colloquial usage. When one encounters such spellings “the convention violated is one of the eyes, and not of the ear”. Furthermore, eye dialect would be used by writers “not to indicate a genuine difference in pronunciation, but the spelling is a friendly nudge to the reader, a knowing look which establishes a sympathetic sense of superiority between the author and reader as contrasted with the humble speaker of dialect”. While the phrase “the humble speaker of dialect” may smack of prescriptivism to the modern reader, this passage is important, as it finally gives a term for a device that has been used in literature for centuries. Krapp was referring to spellings likeenufffor ‘enough’,wimminfor ‘women’,animulzfor ‘animals’ and numerous other examples in which the standard spelling of the word belies in some way its pronunciation. One may envisage these spellings as a sort of insinuation on the part of the author that the character whose speech is depicted so would spell these words in this way, hence demonstrating a level of education and literacy substantially lower than the average

    Collocate networks in the language of crime journalism

    No full text
    Standard procedures for the treatment of collocates, which involve the elaboration of lists of collocates on a two-by-two basis, are far from optimum for the study of connectivity, i.e. observing whether these collocates in turn display a tendency to co-occur or not. This paper explores an alternative strategy that has garnered considerable interest in recent years: that of using Social Network Analysis procedures. Lists of collocates (concgrams) were extracted from a one million word corpus of crime journalism using standard techniques. Gephi software was then used to transform the list of collocates into a network. A small number of collocate pairs were seen to be isolates, i.e. collocating only with each other, while the majority belonged to the giant component, composed of pairs in which at least one member collocates with at least one other word. Modules (clusters of highly interconnected collocates) were identified; these were seen to pertain to specific subject areas. The corpus was then re-examined to see where these clusters of collocates occurred, and co-occurred, and to gauge how much this technique may tell us about the 'aboutness' of particular texts

    Words (don't come easy): The automatic retrieval and analysis of popular song lyrics

    No full text
    The current work will describe the compilation of a large (10 million tokens) corpus of popular song lyrics in English divided into sub-genres: the Sassari Lyrics (SLY) Corpus. The texts were gathered by web crawling the index pages of an online song repository. It will then analyze the keywords of each sub-genre and shared keywords, highlighting similarities and differences between sub-genres. The first part of this paper will discuss the procedures adopted to retrieve the song lyrics, along with metadata such as date, author, album and sub-genre. The repository proved somewhat unreliable regarding the attribution of artists to musical sub-genres, therefore alternative semi-automatic processes had to be developed. Several other reliability issues will be discussed, for example, songs in foreign languages, covers, variation in song titles and artist names are all factors that had to be filtered out or normalized. The second part will present preliminary results concerning the analysis of keywords. While each sub-genre (ALTERNATIVE ROCK, COUNTRY, HIP HOP, HEAVY METAL, POP, RNB and ROCK) had a considerable number of keywords, we noticed that those of some sub-genres, such as HIP HOP and HEAVY METAL, were highly characteristic lexical items, those of others, such as POP and RNB were mainly grammatical items with very high frequencies. The latter two sub-genres share so many keywords that it could be argued that, at least on a textual basis, they are essentially not discernible
    corecore