1,721,203 research outputs found
Eye dialect: translating the untranslatable
The term ‘eye dialect’ was first coined in 1925 by George P. Krapp inThe
English Language in America(McArthur 1998). The term was used to describe
the phenomenon of unconventional spelling used to reproduce colloquial
usage. When one encounters such spellings “the convention violated is one
of the eyes, and not of the ear”. Furthermore, eye dialect would be used by
writers “not to indicate a genuine difference in pronunciation, but the
spelling is a friendly nudge to the reader, a knowing look which establishes a
sympathetic sense of superiority between the author and reader as
contrasted with the humble speaker of dialect”. While the phrase “the
humble speaker of dialect” may smack of prescriptivism to the modern
reader, this passage is important, as it finally gives a term for a device that
has been used in literature for centuries. Krapp was referring to spellings
likeenufffor ‘enough’,wimminfor ‘women’,animulzfor ‘animals’ and
numerous other examples in which the standard spelling of the word belies
in some way its pronunciation. One may envisage these spellings as a sort of
insinuation on the part of the author that the character whose speech is
depicted so would spell these words in this way, hence demonstrating a
level of education and literacy substantially lower than the average
Collocate networks in the language of crime journalism
Standard procedures for the treatment of collocates, which involve the elaboration of lists of collocates on a two-by-two basis, are far from optimum for the study of connectivity, i.e. observing whether these collocates in turn display a tendency to co-occur or not. This paper explores an alternative strategy that has garnered considerable interest in recent years: that of using Social Network Analysis procedures. Lists of collocates (concgrams) were extracted from a one million word corpus of crime journalism using standard techniques. Gephi software was then used to transform the list of collocates into a network. A small number of collocate pairs were seen to be isolates, i.e. collocating only with each other, while the majority belonged to the giant component, composed of pairs in which at least one member collocates with at least one other word. Modules (clusters of highly interconnected collocates) were identified; these were seen to pertain to specific subject areas. The corpus was then re-examined to see where these clusters of collocates occurred, and co-occurred, and to gauge how much this technique may tell us about the 'aboutness' of particular texts
Social Network Analysis and the Analysis of Collocations in the Language of Travel Journalism
Words (don't come easy): The automatic retrieval and analysis of popular song lyrics
The current work will describe the compilation of a large (10 million tokens) corpus of popular song lyrics in English divided into sub-genres: the Sassari Lyrics (SLY) Corpus. The texts were gathered by web crawling the index pages of an online song repository. It will then analyze the keywords of each sub-genre and shared keywords, highlighting similarities and differences between sub-genres. The first part of this paper will discuss the procedures adopted to retrieve the song lyrics, along with metadata such as date, author, album and sub-genre. The repository proved somewhat unreliable regarding the attribution of artists to musical sub-genres, therefore alternative semi-automatic processes had to be developed. Several other reliability issues will be discussed, for example, songs in foreign languages, covers, variation in song titles and artist names are all factors that had to be filtered out or normalized. The second part will present preliminary results concerning the analysis of keywords. While each sub-genre (ALTERNATIVE ROCK, COUNTRY, HIP HOP, HEAVY METAL, POP, RNB and ROCK) had a considerable number of keywords, we noticed that those of some sub-genres, such as HIP HOP and HEAVY METAL, were highly characteristic lexical items, those of others, such as POP and RNB were mainly grammatical items with very high frequencies. The latter two sub-genres share so many keywords that it could be argued that, at least on a textual basis, they are essentially not discernible
The Praat Programme in the L2 Pronunciation Class: Formant Plotting and the Use of Z Scores
- …
