1,721,092 research outputs found
Regularization Approaches for Synthesizing HRTF Directivity Patterns
As an alternative to traditional artificial heads, it is possible to synthesize individual head-related transfer functions (HRTFs) using a so-called virtual artificial head (VAH), consisting of a microphone array with an appropriate topology and filter coefficients optimized using a narrowband least squares cost function. The resulting spatial directivity pattern of such a VAH is known to be sensitive to small deviations of the assumed microphone characteristics, e.g., gain, phase and/or the positions of the microphones. In many beamformer design procedures, this sensitivity is reduced by imposing a white noise gain (WNG) constraint on the filter coefficients for a single desired look direction. In this paper, this constraint is shown to be inappropriate for regularizing the HRTF synthesis with multiple desired directions and three alternative different regularization approaches are proposed and evaluated. In the first approach, the measured deviations of the microphone characteristics are taken into account in the filter design. In the second approach, the filter coefficients are regularized using the mean WNG for all directions. The third approach additionally takes into account several frequency bins into both the optimization and the regularization. The different proposed regularization approaches are compared using analytic and measured transfer functions, including random deviations. Experimental results show that the approach using multiple frequency bands mimicking the spectral resolution of the human auditory system yields the best robustness among the considered regularization approaches
Evaluation of head-tracked binaural auralizations of speech signals generated with a virtual artificial head in anechoic and classroom environments
In order to realize binaural auralizations with head tracking, BRIRs of individual listeners are needed for different head orientations. In this contribution, a filter-and-sum beamformer, referred to as virtual artificial head (VAH), was used to synthesize the BRIRs. To this end, room impulse responses were first measured with a VAH, using a planar microphone array with 24 microphones, for one fixed orientation, in an anechoic and a reverberant room. Then, individual spectral weights for 185 orientations of the listener’s head were calculated with different parameter sets. Parameters included the number and the direction of the sources considered in the calculation of spectral weights as well as the required minimum mean white noise gain (WNGm). For both acoustical environments, the quality of the resulting synthesized BRIRs was assessed perceptually in head-tracked auralizations, in direct comparison to real loudspeaker playback in the room. Results showed that both rooms could be auralized with the VAH for speech signals in a perceptually convincing manner, by employing spectral weights calculated with 72 source directions from the horizontal plane. In addition, low resulting WNGm values should be avoided. Furthermore, in the dynamic binaural auralization with speech signals in this study, individual BRIRs seemed to offer no advantage over non-individual BRIRs, confirming previous results that were obtained with simulated BRIRs
Dynamic Binaural Rendering: The Advantage of Virtual Artificial Heads over Conventional Ones for Localization with Speech Signals
As an alternative to conventional artificial heads, a virtual artificial head (VAH), i.e., a microphone array-based filter-and-sum beamformer, can be used to create binaural renderings of spatial sound fields. In contrast to conventional artificial heads, a VAH enables one to individualize the binaural renderings and to incorporate head tracking. This can be achieved by applying complex-valued spectral weights—calculated using individual head related transfer functions (HRTFs) for each listener and for different head orientations—to the microphone signals of the VAH. In this study, these spectral weights were applied to measured room impulse responses in an anechoic room to synthesize individual binaural room impulse responses (BRIRs). In the first part of the paper, the results of localizing virtual sources generated with individually synthesized BRIRs and measured BRIRs using a conventional artificial head, for different head orientations, were assessed in comparison with real sources. Convincing localization performances could be achieved for virtual sources generated with both individually synthesized and measured non-individual BRIRs with respect to azimuth and externalization. In the second part of the paper, the results of localizing virtual sources were compared in two listening tests, with and without head tracking. The positive effect of head tracking on the virtual source localization performance confirmed a major advantage of the VAH over conventional artificial heads
Windgeräuschanalyse, -synthese und -reduzierung zur Sprachverbesserung mit kompakten Mikrofonarrays
Wind noise refers to random fluctuations of air pressure induced by a wind stream and captured by microphones. In contrast to acoustic noise, wind noise is not generated by propagating sound waves but from pressure fluctuations caused by turbulent air motion. Turbulence results in low-frequency rumbling distortions in audio recordings. Such artifacts are unwanted as they can significantly degrade desired acoustic signals in terms of quality and, in the case of speech, intelligibility. In addition, wind noise can represent a danger for hearing aid users since it might mask safety-critical sounds (ambulance, alarm) and reduce spatial cues. Mechanical solutions to counteract wind noise have been designed for large capsule microphones, i.e., so-called windscreens can diffuse and redirect the turbulence by covering the microphone with a foam hood. However, windscreens are unsuitable for compact devices like smartphones, wearables, action cameras, and hearing aids, as they reduce their usability and portability. For this reason, digital signal processing techniques are preferred to reduce wind noise and enhance the desired signal. Standard noise reduction techniques assume a stationary noise floor that varies slower than, e.g., speech. Hence, the noise statistics are commonly estimated during speech absence by exploiting voice activity detection or speech presence probability. However, wind noise is highly unpredictable and heavily non-stationary due to abrupt wind intensity changes, so that common noise reduction techniques yield unsatisfactory performance. Therefore, wind noise reduction methods have been developed in the past decades based on contrasting temporal and spectral characteristics of acoustic and air turbulence-induced signals. In addition, the miniaturization and the integration of multiple microphones in audio devices enabled the design of multi-channel processing methods to attenuate wind noise. Such methods can combine spatial and spectral enhancement and provide increased performance compared to single-channel methods. Most multi-channel wind noise reduction approaches are derived assuming that the wind noise is spatially uncorrelated, and exploit the coherent nature of propagating acoustic waves to estimate the desired signal and wind noise statistics. This assumption can, however, be violated when microphone arrays with sufficiently small inter-microphone distances are employed and could therefore limit the reduction performance. In this thesis, we present novel contributions to the analysis, synthesis, and reduction of wind noise measured with closely spaced microphones. Furthermore, we propose methods for estimating the wind speed and direction using a compact microphone array. In particular, we show how and under which conditions multi-channel wind noise can be correlated. Similar to acoustic fields that exhibit a spatially diffuse coherence model, we approximate the spatial coherence of wind noise contributions with a semi-empirical model, namely the Corcos model. According to the Corcos model, wind noise propagates across space at nearly the wind speed and in the same direction. However, the wind noise coherence decreases anisotropically, i.e., at different rates along the streamwise (parallel to the wind direction) and the spanwise direction (orthogonal to the wind direction). We validate the Corcos model with wind noise data collected indoors (wind tunnel) and outdoors (atmospheric wind). In addition, we show that specific temporal and spectral characteristics are related to the flow speed. As a first main contribution, we combine the abovementioned observations to design a multi-channel wind noise generation framework. The synthetic generation of noise samples facilitates developing and evaluating noise reduction techniques in a controlled environment. In fact, wind noise is generally difficult to isolate from outdoor recordings, where different acoustic sources can be simultaneously active. Furthermore, it reduces to a large extent the time necessary to collect a sufficient amount of data. The proposed generation method takes wind speed and direction profiles as input and yields synthetic wind noise signals exhibiting a spatial coherence according to the Corcos model. In addition, we propose a multi-channel wind noise detector based on the contrasting spatial characteristics of speech and wind noise, and three methods to enhance the simulation of acoustic or turbulence-induced signals with a predefined spatial coherence. The second main contribution is the design of two methods for estimating either or both the wind speed and direction based on microphone signals measured with a compact array. Although conventional instruments can achieve high accuracy, employing closely spaced microphones has many advantages, e.g., integrability, high scalability, and a cost-contained implementation. As the Corcos model depends on the wind velocity, the spatial coherence of wind noise provides information on the sought quantities. We first present a deep learning-based method to infer the wind speed consisting of a feedforward neural network trained with synthetic wind noise correlated according to the Corcos model. Pair-wise spatial coherence functions are used as an input feature to obtain a wind speed estimate as output. We then present a method to resolve the speed and direction by fitting the measured spatial coherence to the analytical Corcos model in the least-squares sense. The methods are evaluated using data collected indoors and outdoors and labeled with an ultrasonic anemometer.As a third main contribution, we propose a multi-channel wind noise reduction method for closely spaced microphones based on a parametric multi-channel Wiener filter. In contrast to the well-established assumption of uncorrelated wind noise in multi-channel recordings, we assume a non-zero spatial coherence at low frequencies. For this reason, we recursively estimate the spatial coherence matrix based on the noisy microphone observations through an expectation-maximization approach. In addition, we prove the equivalence of two recently developed noise power spectral density estimation methods when uncorrelated wind noise is assumed. We propose an approximation of both estimators, which is independent of the speech propagation vector under specific conditions. A frequency-dependent wind noise indicator, namely the difference-to-sum power ratio, is used as a parameter that trades off speech distortion and noise reduction in a parametric Wiener postfilter. An evaluation of improvements in speech quality, signal-to-noise ratio, and intelligibility is carried out using both simulated and measured wind noise samples in comparison to an existing method. Finally, we present a method to reduce wind noise in B-format signals that preserves the spatial distribution of the noise field in the surround format after the reduction step. An omnidirectional-to-dipole power ratio is formulated and employed as a trade-off between the desired signal distortion and noise reduction.Windgeräusche bezeichnen zufällige Schwankungen des Luftdrucks aufgrund eines Luftstromes, welche von Mikrofonen erfasst werden können. Im Gegensatz zu akustischem Lärm werden Windgeräusche nicht durch sich ausbreitende Schallwellen erzeugt, sondern durch Druckschwankungen, welche durch turbulente Luftbewegungen verursacht werden. Die Turbulenzen führen zu niederfrequenten, rumpelartigen Verzerrungen in Audioaufnahmen. Solche Artefakte sind unerwünscht, da sie die Qualität der gewünschten akustischen Signale und, im Falle von Sprache, die Verständlichkeit erheblich beeinträchtigen können. Darüber hinaus können Windgeräusche eine Gefahr für Hörgeräteträger darstellen, da sie sicherheitskritische Geräusche (Krankenwagen, Alarm) überdecken und räumliche akustische Merkmale verringern können. Für große Kapselmikrofone wurden mechanische Lösungen entwickelt, um Windgeräusche zu reduzieren. Sogenannte Windschütze, schaumstoffhaubenartige Abdeckungen des Mikrofons, können Turbulenzen zerstreuen und umlenken. Windschütze sind jedoch für die Verwendung auf kompakten Geräten wie Smartphones, Wearables, Action-Kameras und Hörgeräten ungeeignet, da sie deren Nutzbarkeit und Tragbarkeit einschränken. Aus diesem Grund werden Verfahren der digitalen Signalverarbeitung bevorzugt, um Windgeräusche zu unterdrücken und das gewünschte Signal zu verbessern. Standardverfahren zur Rauschunterdrückung gehen von einem stationären Grundrauschen aus, welches sich langsamer verändert als beispielsweise Sprache. Daher werden die Geräuschstatistiken des Rauschens üblicherweise während der Abwesenheit von Sprache geschätzt, unter Zuhilfenahme einer Sprachaktivitätsdetektion oder der Wahrscheinlichkeit der Anwesenheit von Sprache. Windgeräusche, hingegen, sind in hohem Maße unberechenbar , da abrupte Änderungen der Windintensität starke Instationaritäten bedingen, so dass herkömmliche Verfahren zur Geräuschreduzierung keine zufriedenstellende Leistung erbringen. Daher wurden in den letzten Jahrzehnten Methoden zur Reduzierung von Windgeräuschen entwickelt, die auf den gegensätzlichen zeitlichen und spektralen Eigenschaften von akustischen und aerodynamischen Signalen basieren. Darüber hinaus ermöglichten die Miniaturisierung und die Integration mehrerer Mikrofone in Audiogeräte die Entwicklung von mehrkanaligen Verarbeitungsmethoden zur Dämpfung von Windgeräuschen. Solche Methoden können räumliche und spektrale Verstärkung kombinieren und erbringen im Vergleich zu einkanaligen Methoden eine bessere Leistung. Die meisten mehrkanaligen Ansätze zur Windgeräuschreduzierung gehen von der Annahme aus, dass das Windgeräusch räumlich unkorreliert ist und nutzen die kohärente Natur der sich ausbreitenden akustischen Wellen, um die gewünschten Signal- und Windgeräuschstatistiken zu schätzen. Diese Annahme kann jedoch verletzt werden, wenn Mikrofonanordnungen mit ausreichend kleinen Abständen zwischen den Mikrofonen verwendet werden, wodurch die Leistung der Geräuschreduzierung reduziert wird. In dieser Arbeit stellen wir neue Beiträge zur Analyse, Synthese und Reduzierung von Windgeräuschen vor, die mit eng beieinander liegenden Mikrofonen gemessen wurden. Darüberhinaus schlagen wir Methoden zur Schätzung von Windgeschwindigkeit und -richtung mit Hilfe eines kompakten Mikrofonarrays vor. Insbesondere zeigen wir, wie und unter welchen Bedingungen mehrkanalige Windgeräusche korreliert sind. Ähnlich wie bei akustischen Feldern, die ein räumlich diffuses Kohärenzmodell aufweisen, approximieren wir die räumliche Kohärenz von Windgeräuschbeiträgen mit einem semi-empirischen Modell, dem Corcos-Modell. Dem Corcos-Modell folgend breitet sich das Windgeräusch mit nahezu gleicher Windgeschwindigkeit und Richtung im Raum aus. Es verliert jedoch anisotrop an Kohärenz, d.h. mit unterschiedlichen Raten entlang der Strömungsrichtung (parallel zur Windrichtung) und der Spannweitenrichtung (orthogonal zur Windrichtung). Wir validieren das Corcos-Modell mit Windgeräuschdaten, die in Innenräumen (Windkanal) und im Freien (atmosphärischer Wind) gemessen wurden. Darüber hinaus zeigen wir, dass bestimmte zeitliche und spektrale Windmerkmale mit der Strömungsgeschwindigkeit zusammenhängen. Als ersten Hauptbeitrag kombinieren wir die oben genannten Beobachtungen, um ein neuartiges mehrkanaliges System zur Erzeugung von Windgeräuschen zu entwickeln. Die vorgeschlagene synthetische Erzeugung von Geräuschproben ermöglicht dabei die Entwicklung und Bewertung von Techniken zur Geräuschreduzierung in einer kontrollierten Umgebung. Windgeräusche lassen sich in der Regel nur schwer von Außenaufnahmen isolieren, bei denen verschiedene akustische Quellen gleichzeitig aktiv sein können. Außerdem wird dadurch die Zeit, die zur Erfassung einer ausreichenden Menge an Daten erforderlich ist erheblich verkürzt. Die vorgeschlagene Generierungsmethode nimmt Windgeschwindigkeits- und -richtungsprofile als Eingabe und erzeugt synthetische Windgeräuschsignale, die eine räumliche Kohärenz gemäß des Corcos-Modells aufweisen. Darüber hinaus schlagen wir einen mehrkanaligen Windgeräuschdetektor vor, der auf den unterschiedlichen räumlichen Eigenschaften von Sprache und Windgeräuschen basiert, sowie drei Methoden zur Verbesserung der Simulation akustischer Signale mit einer vordefinierten räumlichen Kohärenz. Der zweite Hauptbeitrag ist die Entwicklung von zwei Methoden zur Schätzung der Windgeschwindigkeit und -richtung auf der Grundlage von Mikrofonsignalen, welche mit einem kompakten Array gemessen wurden. Obwohl mit herkömmlichen Instrumenten eine hohe Genauigkeit erreicht werden kann, hat die Verwendung von Mikrofonen mit geringem Abstand viele Vorteile, z.B. Integrierbarkeit, hohe Skalierbarkeit und eine kostengünstige Implementierung. Da Werte im Corcos-Modell von der Windgeschwindigkeit abhängen, liefert die räumliche Kohärenz von Windgeräuschen Informationen über die gesuchten Größen. Wir stellen zunächst eine auf Deep Learning basierte Methode zur Schätzung der Windgeschwindigkeit vor, die aus einem feedforward neuronalen Netz besteht, das mit synthetischen Windgeräuschen trainiert wurde, welche dem Corcos-Modell entsprechend korreliert sind. Paarweise räumliche Kohärenzfunktionen werden als Eingangsmerkmal verwendet, um eine Windgeschwindigkeitsschätzung als Ergebnis zu erhalten. Anschließend stellen wir eine Methode zur Schätzungvon Geschwindigkeit und Richtung vor, indem wir die gemessene räumliche Kohärenz an das analytische Corcos-Modell im Sinne des kleinsten quadratischen Fehlers annähern. Die Methoden werden anhand von Daten evaluiert, die in Innenräumen und im Freien gesammelt und mithilfe eines Ultraschallanemometer annotiert wurden.Als dritten Hauptbeitrag schlagen wir eine mehrkanalige Methode zur Reduzierung von Windgeräuschen für eng beieinander liegende Mikrofone vor. Im Gegensatz zu der üblichen Annahme unkorrelierter Windgeräusche bei mehrkanaligen Aufnahmen gehen wir von einer räumlichen Kohärenz ungleich Null bei niedrigen Frequenzen aus. Aus diesem Grund schätzen wir die räumliche Kohärenzmatrix rekursiv auf der Grundlage der Mikrofonbeobachtungen durch einen Erwartungsmaximierungsansatz. Darüber hinaus beweisen wir die Gleichwertigkeit zweier kürzlich entwickelter Methoden zur Schätzung der spektralen Rauschleistungsdichte, wenn unkorrelierte Windgeräusche angenommen werden. Wir schlagen eine Approximation beider Schätzer vor, die unter bestimmten Bedingungen unabhängig vom Sprachausbreitungsvektor ist. Eine Bewertung der Verbesserungen bei der Sprachqualität, dem Signal-Rausch-Verhältnis und der Verständlichkeit wird anhand von simulierten und gemessenen Windgeräuschproben im Vergleich zu einer bestehenden Methode durchgeführt. Schließlich stellen wir eine Methode zur Reduzierung von Windgeräuschen in Signalen im B-Format vor, bei der die räumliche Verteilung des Geräuschfeldes im Surround-Format nach dem Reduzierungsschritt erhalten bleibt. Es wird ein Verhältnis von omnidirektionaler zu dipolarer Leistung formuliert und als Kompromiss zwischen der gewünschten Signalverzerrung und der Rauschreduzierung verwendet
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Smoothing individual head-related transfer functions in the frequency and spatial domains
When re-synthesizing individual head related transfer functions (HRTFs) with a microphone array, smoothing HRTFs spectrally and/or spatially prior to the computation of appropriate microphone filters may improve the synthesis accuracy. In this study, the limits of the associated HRTF modifications, until which no perceptual degradations occur, are explored. First, complex spectral smoothing of HRTFs into constant relative bandwidths was considered. As a prerequisite to complex smoothing, the HRTF phase spectra were substituted by linear phases, either for the whole frequency range or above a certain cut-off frequency only. The results indicate that a broadband phase linearization of HRTFs can be perceived for certain directions/subjects and that the thresholds can be predicted by a simple model. HRTF phase spectra can be linearized above 1 kHz without being detectable. After substituting the original phase by a linear phase above 5 kHz, HRTFs may be smoothed complexly into constant relative bandwidths of 1/5 octave, without introducing noticeable artifacts. Second, spatially smoother HRTF directivity patterns were obtained by levelling out spatial notches. It turned out that spatial notches do not have to be retained if they are less than 29 dB below the maximum level in the directivity pattern. (C) 2014 Acoustical Society of America
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
