1,720,978 research outputs found
Modeling the perception of children's age from speech acoustics
Full text access from Treasures at UT Dallas is restricted to current UTD affiliates.Adult listeners were presented with /hVd/ syllables spoken by boys and girls ranging from 5 to 18 years of age. Half of the listeners were informed of the sex of the speaker; the other half were not. Results indicate that veridical age in children can be predicted accurately based on the acoustic characteristics of the talker's voice and that listener behavior is highly predictable on the basis of speech acoustics. Furthermore, listeners appear to incorporate assumptions about talker sex into their estimates of talker age, even when information about the talker's sex is not explicitly provided for them. © 2018 Acoustical Society of America.National Science Foundation (Grant No. 1124479)School of Behavioral and Brain Science
Perceiving foreign-accented speech with decreased spectral resolution in single- and multiple-talker conditions
To determine the effect of reduced spectral resolution on the intelligibility of foreign-accented speech, vocoder-processed sentences from native and Mandarin-accented English talkers were presented to listeners in single- and multiple-talker conditions. Reduced spectral resolution had little effect on native speech but lowered performance for foreign-accented speech, with a further decrease in multiple- talker conditions. Following the initial exposure, foreign-accented speech with reduced spectral resolution was less intelligible than unprocessed speech in both single- and multiple-talker conditions. Intelligibility improved with extended exposure, but only for single- talker conditions. Results indicate a perceptual impairment when perceiving foreign-accented speech with reduced spectral resolution.School of Behavioral and Brain Science
Voice gender and the segregation of competing talkers: Perceptual learning in cochlear implant simulations
Full text access from Treasures at UT Dallas is available only to current UTD affiliates.Two experiments explored the role of differences in voice gender in the recognition of speech masked by a competing talker in cochlear implant simulations. Experiment 1 confirmed that listeners with normal hearing receive little benefit from differences in voice gender between a target and masker sentence in four- and eight-channel simulations, consistent with previous findings that cochlear implants deliver an impoverished representation of the cues for voice gender. However, gender differences led to small but significant improvements in word recognition with 16 and 32 channels. Experiment 2 assessed the benefits of perceptual training on the use of voice gender cues in an eight-channel simulation. Listeners were assigned to one of four groups: (1) word recognition training with target and masker differing in gender; (2) word recognition training with same-gender target and masker; (3) gender recognition training; or (4) control with no training. Significant improvements in word recognition were observed from pre- to post-test sessions for all three training groups compared to the control group. These improvements were maintained at the late session (one week following the last training session) for all three groups. There was an overall improvement in masked word recognition performance provided by gender mismatch following training, but the amount of benefit did not differ as a function of the type of training. The training effects observed here are consistent with a form of rapid perceptual learning that contributes to the segregation of competing voices but does not specifically enhance the benefits provided by voice gender cues.National Institute of Deafness and other Communication Disorders F31 DC9537; National Science Foundation grant 1124479.School of Behavioral and Brain Science
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Production and Perception of Affective Prosody by Adults with Autism Spectrum Disorder
Affective prosody, defined as the use of paralinguistic elements in speech to convey emotion, is important for effective social functioning. While generally a trivial task for typically-developing (TD) adults, individuals with autism spectrum disorder (ASD) present with significant challenges in social communication and interaction, including prosody. Previous research has shown that talkers with ASD produce pragmatic prosody with increased variability in fundamental frequency (f0, which is closely correlated with voice pitch), but it was unclear whether those differences carry over to speaking tasks involving emotion elicitation. A controlled set of expressive speech recordings was obtained from talkers with ASD and controls in five emotion contexts: angry, happy, interested, sad, and neutral. Emotion-specific group differences in f0, intensity, and duration were found in multiple speech types, and the pattern of results was characterized by inconsistent and exaggerated patterns of affective prosody production in talkers with ASD compared to controls. The perceptual relevance of the affective acoustic differences was tested in three listening experiments involving talkers and listeners with ASD and controls. The first two experiments involving TD listeners were designed to examine the perceptual impact of increased f0 variability and intensity found in recordings produced by talkers with ASD. Compared to the intensity manipulation, modifying the f0 contour had a larger impact on emotion recognition accuracy. The third experiment was designed to compare perception of affective prosody in listeners with ASD and controls using unmodified stimuli, and revealed that differences in affective prosody perception were more closely related to talker group production differences than listener group differences. The results are consistent with previous work in face perception showing increased emotion identification rates but lower naturalness ratings when listeners responded to stimuli produced by individuals with ASD. The findings are interpreted within the context of the speech attunement framework, which suggests that individuals with ASD lack the motivation to attune their prosodic speech to sound like TD talkers
Perception of Novel Sounds in the Presence of Background Noise : Comparison Between Individuals with Normal Hearing, Cochlear Implant Users, and Recurrent Neural Networks
The goal of this dissertation is to investigate how listeners and learning machines cope with the
ambiguity caused by interfering multiple novel sound sources. Starting from an ambiguous
auditory scene with competing sound sources, this dissertation investigates how a particular
sound source draws listeners’ attention while the remaining sources lose their salience and
become background (noise). Listeners’ perception of competing novel sounds is investigated in a
series of experiments that varied in terms of listening conditions, simulating the difficulties
experienced by hearing-impaired individuals in noise. In Chapter 1, the mechanisms behind
listeners' perception of speech in the presence of competing sounds are reviewed. Chapter 2
describes three experiments that investigated the recognition of novel sounds in the presence of
background noise. The chapter begins with a replication of a previous study, providing evidence
that listeners can segregate a novel target sound from the competing distractor only if it repeats
across different distractors. A subsequent experiment tested the hypothesis that listeners’ ability
to detect change in a sound depends on their knowledge of its source, which is gained via
repetition. It is concluded that listeners are able to perceptually learn patterns of the repeating
target while suppressing the changes in the masker stream. Two neural network architectures
previously employed to study mechanisms of learning, generalized Hebbian and anti-Hebbian,
are evaluated. It is shown that the generalized Hebbian learning network produces similar results
to those obtained from the listeners. Experiments in Chapter 3 provide evidence that recognition
of a novel target sound becomes robust against new (unheard) distractors when listeners go
through an exposure stage in which the target is presented repeatedly across multiple distractors.
Chapter 3 concludes by reporting experiments 3-2 and 3-3 that investigated recognition of
consonant-vowel-consonant-vowel (CVCV) words in the presence of novel distractors.
Experiment 3-2 showed that upon exposing the listeners to target tokens across multiple
distractors, the process of learning new CVCV tokens shifts from context-specificity to an
adaptation-plus-prototype mechanism. The goal in experiment 3-3 was to investigate whether or
not cochlear implant users, who have limited spectral resolution, would show the same behavior
as listeners with normal hearing in experiment 3-2. The main goal in Chapter 4 is to investigate
the extent to which the findings in experiment 3-2 can be replicated by recurrent neural networks
(RNNs). This chapter begins with a brief introduction to RNNs and long short-term memories
(LSTMs). In experiment 4-1 a recurrent LSTM auto-encoder was trained to reconstruct an input
CVCV target when mixed with a distractor with or without the presence of a context sequence
prior to the input. It was shown that the network could reconstruct the input with better accuracy
when the context sequence contained the repeating CVCV target across multiple distractors.
Furthermore, similar to the findings in experiment 3-2, the presence of such a context sequence
improved the network’s generalizability to unseen data (novel distractors). Experiment 4-2
showed that the presence of the context sequence led to an improved semi-supervised speech
enhancement algorithm that recovered the target CVCV tokens while suppressing the distractors
- …
