1,720,966 research outputs found
The effects of recording devices and software on phonetic analysis
Note: This talk has not gone through a process of peer review, and findings should therefore be treated as preliminary and subject to change.
SOAS Linguistics Webinars: The effects of recording devices and software on phonetic analysis
7 July 2021
Dr Chelsea Sanker
Yale University
Abstract: Because of restrictions on in-person research due to Covid-19, researchers are now relying on remotely recorded data to a much greater extent than has been typical in the past. Given the change in methodology, it is important to know how remote recording might affect acoustic measurements, either because of different recording devices used by participants and consultants recording themselves or because of video-conferencing software used to make interactive recordings. This study investigates audio signal fidelity across different in-person recording equipment and remote recording software when compared to a solid-state digital audio recording device. We show that the choice of equipment and software can have a large effect on acoustic measurements, including measurements of frequency, duration, and noise. The issues do not just reflect decreased reliability of measurements; some measurements are systematically shifted in particular recording conditions. These results show the importance of carefully considering and documenting equipment choices. In particular, any cross-linguistic or cross-speaker comparison needs to account for possible effects of differences in which devices or software platforms were used.
Based on joint work by Chelsea Sanker, Sarah Babinski, Roslyn Burns, Marisha Evans, Jeremy Johns, Juhyae Kim, Slater Smith, Natalie Weber, and Claire Bowern (all Yale University)
How to cite: Sanker, Chelsea. 2021. The effects of recording devices and software on phonetic analysis. Talk presented at SOAS Linguistics Webinar 2021-07-07.
Also available on YouTube: https://youtu.be/ZZ8lYsR9Ub
Comparison of Phonetic Convergence in Multiple Measures
During interaction, speakers acquire characteristics more similar to characteristics of their interlocutors’ speech and other behaviors; this is known as convergence. In addition to non-linguistic characteristics such as posture (Dijksterhuis and Bargh 2001) and fidgeting movements (Chartrand and Bargh 1999), this has been found in many characteristics of speech, including pitch (Babel and Bulatov 2011), vowel formants (Babel 2012), intensity (Gregory and Hoyt 1982), lexical items (Ireland et al. 2011), syntactic constructions (Nilsenová and Noltig 2010), and timing of conversational turns and pauses (Street 1984).
A variety of explanations for convergence have been proposed. The main explanations for linguistic convergence are understanding-based (e.g. Street and Giles 1982), socially motivated (e.g. Eckert 2001), or automatic (Dijksterhuis and Bargh 2001). Each of these explanations has elements of support from a range of experiments. As more studies add details to the range of influences on convergence and the characteristics affected by it, we can make a clearer picture of the cognitive representations of linguistic characteristics and their dynamics.
The amount of phonetic convergence between speakers has been associated with several factors, such as race (Babel 2012), gender (Pardo 2010), age (Labov 2006), nationality (Giles, Coupland, and Coupland 1991), native language (Kim et al. 2011), and interlocutor status (Gregory and Webster 1996, Bane et al. 2010/2014). One large factor correlated with convergence is positiveness of interlocutors’ opinions of each other, using a range of positiveness measures. Convergence occurs to a greater degree between people with positive relationships (Bernieri and Rosenthal 1991) and greater convergence leads to more positive opinions of interlocutors (Giles et al. 1973). Convergence also depends on characteristics of the individuals involved: people who are more concerned with social status exhibit more convergence (Pardo 2006, Natale 1975).
However, the extent to which convergence in different characteristics aligns remains unclear; correlation between convergence in different measures has not been part of many previous studies, in part due to the different experimental designs suited to measuring different characteristics. Convergence in word-level or phoneme-level characteristics such as vowel formants (e.g. Babel 2012) and voice onset time (e.g. Nielsen 2011) or in overall perceived similarity (e.g. Goldinger 1998) have often been featured in shadowing tasks, in which the participants were recorded repeating words immediately after hearing a recording of them (e.g. Goldinger 1998, Babel 2012) and in interactive tasks designed to elicit repetition of words or concepts (e.g. map task in Pardo 2006; mazes in Garrod and Doherty 1994). Larger-scale patterns such as lexical choice, syntax, intensity, turn durations, and pause durations have often been measured in natural conversational settings (e.g. Gregory and Hoyt 1982, Natale 1975, Bernieri and Rosenthal 1991).
While convergence has been observed in each of these measures and some of them are independently correlated with some of the same social conditions, e.g. ratings of closeness and amount in common (Pardo 2012) and absolute measurements, e.g. number of turns and amount of overlapping speech (Levitan et al. 2012), it is not clear that convergence in each of these characteristics behaves the same way, because of the lack of correlation tests between measures. The results of one recent study by Pardo et al. (2015) on F1, F2, and F0 suggest that the pattern of convergence can differ depending on which measure is used. Degree of convergence based on perceptual testing of similarity has been correlated with the convergence exhibited in characteristics such as pitch and vowel duration (e.g. Pardo 2010), though in some studies researchers have not found correlation between the results of the perceptual test and the phonetic characteristic tested (e.g. Babel and Bulatov 2011, Pardo et al. 2012).
In this study, the relative patterns of convergence were investigated for eight different characteristics across pairs of interlocutors: F1, F2, vowel duration, F0, intensity, turn duration, duration of pauses marking the transition to a new speaker, and duration of pauses after which the same speaker resumed. Convergence was calculated by comparing the average value in each measure for each participant over four time periods and then comparing the average difference between partners’ averages through those time periods. After calculating convergence in each measure, the correlation between convergence in each possible combination of measures was analyzed, to determine how closely convergence in one measure aligns with convergence in other measures.
I hypothesized that convergence in each measure would be positively correlated with the other measure, if the same mechanism underlies convergence in each of these measures; degree of convergence within a pair in vowel formants, F0, intensity, and timing of turns, turn-switching pauses, and within-turn pauses would be predictive of degree of convergence in each of the other measures. The null hypothesis was that convergence in each measure is not correlated with other measures; degree of convergence within a pair in vowel formants, F0, intensity, and timing of turns, turn-switching pauses, and within-turn pauses will not be predictive of degree of correlation in any other measure, suggesting that convergence is mediated through individual differences in attention to different acoustic and timing characteristics of speech, or that there is a slightly different process underlying convergence in different characteristics. Either result has implications for experimental design in convergence research.This working paper is copyrighted, and is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) - see https://creativecommons.org/licenses/by-nc-nd/4.0
Frequency-predicted shifts independent of word-specific phonetic details
Some sound changes seem to proceed at different rates depending on lexical frequency; these are often interpreted as reflecting phonetically detailed exemplar memories, with changes spreading via lexical diffusion (Pierrehumbert 2002; Bybee 2012). However, such patterns do not necessarily require word-specific phonetic details. Variation associated with lexical frequency also exists when there is no evidence for a change in progress, which might be explained by the process of lexical access: Higher lexical frequency facilitates activation, causing faster and more reduced productions (Gahl et al. 2012, Kahn & Arnold 2012, Jurafsky et al. 2002). This work examines how repeated exposure to particular words influences listeners’ category boundary between aspirated and unaspirated stops in those words. Listeners’ VOT category boundary is lowered after exposure to shortened VOT stimuli and also after exposure to lengthened VOT stimuli. These results suggest that frequency-related sound change can largely be explained by frequency directly influencing reduction in phonetic implementation and perceptual access. The size of the effect differed based on the acoustic characteristics of the exposure stimuli; this may suggest a role of word-specific phonetic details, but could also reflect different levels of activation due to the prototypicality of the stimuli
Lexical ambiguity and acoustic distance in discrimination
This work presents a perceptual study on how acoustic details and knowledge of the lexicon influence discrimination decisions. English-speaking listeners were less likely to identify phonologically matching items as the same when they differed in vowel duration, but differences in mean F0 did not have an effect. Although both are components of English contrasts, the results only provide evidence for attention to vowel duration as a potentially contrastive cue. Lexical ambiguity was a predictor of response time. Pairs with matching duration were identified more quickly than pairs with distinct duration, but only among lexically ambiguous items, indicating that lexical ambiguity mediates attention to acoustic detail. Lexical ambiguity also interacted with neighborhood density: Among lexically unambiguous words, the proportion of \u27same\u27 responses decreased with neighborhood density, but there was no effect among lexically ambiguous words. This interaction suggests that evaluating phonological similarity depends more on lexical information when the items are lexically unambiguous
Recommended from our members
The Influence of Lexical Frequency, Phonotactic Probability, and Neighborhood Density on Word Identification
I examine how lexical frequency, neighborhood density, and phonotactic probability predict word identification, as well as the ways that they relate to each other and to the segments within a word. In a two-alternative forced choice task for identification of monosyllabic English words in noise, accuracy was higher for words with higher lexical frequency, lower neighborhood density, and higher phonotactic probability. However, all three characteristics also differ based on the segments within a word. Adding vowel as a predictor of accuracy largely eliminated the observed effects of phonotactic probability and neighborhood density. This relationship with vowel quality might suggest two directions of effects that can influence misperception patterns and subsequently sound change: Phoneme frequency causing directional confusions, or directional confusions changing phoneme frequency
Convergence Doesn\u27t Show Lexically-Specific Phonetic Detail
Can phonetic convergence be lexically specific, providing evidence that representations include word-specific phonetic detail, or does it occur only at a phonological level? Some studies find more convergence in lower frequency words, which is interpreted as evidence for word-specific representations. However, this result has not been consistently replicated, and provides only indirect evidence for word-specific convergence. I more directly test the possibility of word-specific convergence in a shadowing task with different words manipulated in opposing directions; word-specific acoustic details are reflected in immediate repetition, but not in the post-task productions that would indicate shifts in the representation. I also examine a possible alternative source of apparent frequency-conditioned convergence. In a reading task with no exposure to other speakers, frequency was a predictor of speakers becoming more similar to each other in their second reading of a word; because effects of repetition are influenced by lexical frequency, apparent frequency-conditioned convergence can be produced as an artifact of the repetition inherent in shadowing tasks
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
- …
