1,720,960 research outputs found

    Distributed Speech Enhancement in Wireless Acoustic Sensor Networks

    No full text
    In digital speech communication applications like hands-free mobile telephony, hearing aids and human-to-computer communication systems, the recorded speech signals are typically corrupted by background noise. As a result, their quality and intelligibility can get severely degraded. Traditional noise reduction approaches process signals recorded by microphone arrays using centralized beamforming technologies. Recent advances in micro-electro-mechanical systems and wireless communications enable the development of wireless sensor networks (WSNs), where low-cost, low-power and multi-functional wireless sensing devices are connected via wireless links. Compared with conventional localized and regularly arranged microphone arrays, wireless sensor nodes can be randomly placed in environments and thus cover a larger spatial field and yield more information on the observed signals. This thesis explores some problems on multi-microphone speech enhancement for wireless acoustic sensor networks (WASNs), such as distributed noise reduction processing, clock synchronization and privacy preservation. First, we develop a distributed delay-and-sum beamformer (DDSB) for speech enhancement in WASNs. Due to limited power of each wireless device, signal processing algorithms with low computational complexity and low communication cost are preferred in WASNs. Distributed signal processing allows that each node only communicates with its neighboring nodes and performs local processing, where communication load and computational complexity are distributed over all nodes in the network. Without central processor and network topology constraint, the DDSB algorithm estimates the desired speech signal via local processing and local communication. The DDSB algorithm is based on an iterative scheme. More specifically, in each iteration, pairs of neighboring nodes update their estimates according to the principle of traditional delay-and-sum (DSB) beamformer. The estimation of the DDSB converges asymptotically to the optimal solution of the centralized beamformer. However, experimental study indicates that the noise reduction performance of the DDSB is at the expense of a higher communication cost, which can be a serious drawback in practical applications. Therefore, in the second part of this thesis, a clique-based distributed beamformer (CbDB) has been proposed to reduce communication costs of the original DDSB algorithm. In the CbDB, nodes in two neighboring non-overlapping cliques update their estimates simultaneously per iteration. Since each non-overlapping clique consists of multiple nodes, the CbDB allows more nodes to update their estimates and leads to lower communication costs than the original DDSB algorithm. Furthermore, theoretical and experimental studies have shown that the CbDB converges to the centralized beamformer and is more robust for sensor nodes failures in WASNs. In the third part of this thesis, we propose a privacy preserving minimum variance distortionless response (MVDR) beamformer for speech enhancement in WASNs. Different wireless devices in WASNs generally belong to different users. We consider a scenario where a user joins the WASN and estimates his desired source via the WASN, but wants to keep his source of interest private. To introduce a distributed MVDR beamformer in such scenario, a distributed approach is first proposed for recursively estimation of the inverse of the correlation matrix in randomly connected WASNs. This distributed approach is based on the fact that using the Sherman-Morrison formula, estimation of the inverse of the correlation matrix can be seen as a consensus problem. By hiding the steering vector, the privacy preserving MVDR beamformer can reach the same noise reduction performance as its centralized version. In the final part of this thesis, we investigate clock synchronization problems for multi-microphone speech enhancement in WASNs. Each wireless device in WASNs is equipped with an independent clock oscillator, and therefore clock differences are inevitable. However, clock differences between capturing devices will cause signal drift and lead to severe performance degradation of multi-microphone noise reduction algorithms. We provide theoretical analysis of the effect of clock synchronization problems on beamforming technologies and evaluate the use of three different clock synchronization algorithms in the context of multi-microphone noise reduction. Our experimental study shows that the achieved accuracy of the three clock synchronization algorithms enables sufficient accuracy of clock synchronization for the MVDR beamformer in ideal scenarios. However, in practical scenarios with measurement uncertainty or noise, the output of the MVDR beamformer with time-stamp based clock synchronization algorithms gets degraded, while the accuracy of signal based clock synchronization algorithms is still enough for the MVDR beamformer, albeit at a much higher communication cost.Intelligent SystemsElectrical Engineering, Mathematics and Computer Scienc

    Audio-based game for visually impaired children

    No full text
    This thesis was made for the Bachelor Graduation Project (Electrical Engineering). The purpose of the project was to design an audio-based game for visually impaired children. In this thesis the gameplay and the graph- ical user interface are designed. We choose to make a simplified dungeon crawler in the programming language Python. We designed tutori- als and levels for the game. For the tutorials the methods interaural time difference and interaural intensity difference are used to simulate localised audio. For the levels more advanced audio simulation methods are used, but these are provided by two other groups of students. A graphical user interface is made for validation purposes and for parents and caretakers of the visually impaired children. The game is tested with visually impaired children and sighted students. The controls of the game were too complex for young children and the game was not completely accessible for the visually impaired. However, almost all test subjects were able to learn the mechanics of the game and complete levels on their own.Electrical EngineeringElectrical Engineering, Mathematics and Computer Scienc

    Heart rate monitoring using PPG signals

    No full text
    This bachelor’s thesis discusses the idea and implementation of real-time heart rate monitoring using photo- plethysmography (PPG) signals. PPG signals measured from the wrist are often subjected to distortion and noise. Signal processing techniques that tackle these issues were investigated and implemented. The algo- rithm of choice was JOSS. JOSS consists of sparse signal reconstruction, spectral subtraction and spectral peak tracking. The combination of these techniques is what results in a robust heart rate tracking algorithm. An improvement, in terms of computing power, was made by analyzing general properties of PPG signals. The outcome is a new framework named SMART, which combines the HR tracking power of a robust HR monitor- ing system, like JOSS, with the speed of faster HR monitoring methods. The overall result is a fast algorithm that computes the heart rate with an average absolute error of 1.42. This result shows that PPG based heart rate monitoring on wearable devices has great potential for fitness and medical purposes.Electrical Engineering, Mathematics and Computer ScienceMicroelectronicsCircuits and System

    Heart Rate Monitoring Using PPG Signals

    No full text
    Monitoring the heart rate using PPG signals that is worn on the wrist has the downside that it is susceptible to motion. The device on the wrist moves along with the motions of the arm, consequently creating artifacts in the measurement. There are many signal processing techniques that can remove these artifacts. This thesis discusses one of the possible ways to remove motional artifacts from PPG signals, which is the TROIKA framework. This framework consists of three parts: signal decomposition, sparse signal reconstruction and spectral peak tracking. The main subject in this thesis is removal of motional artifacts by signal decomposition. Two methods have been proposed to identify components that belong to motional artifacts. Both methods worked as intended for removing those artifacts. The results of the TROIKA framework that is put together were within expectations.Electrical Engineering, Mathematics and Computer ScienceMicroelectronicsEE3L1

    Advances in DFT-Based Single-Microphone Speech Enhancement

    No full text
    The interest in the field of speech enhancement emerges from the increased usage of digital speech processing applications like mobile telephony, digital hearing aids and human-machine communication systems in our daily life. The trend to make these applications mobile increases the variety of potential sources for quality degradation. Speech enhancement methods can be used to increase the quality of these speech processing devices and make them more robust under noisy conditions. The name "speech enhancement" refers to a large group of methods that are all meant to improve certain quality aspects of these devices. Examples of speech enhancement algorithms are echo control, bandwidth extension, packet loss concealment and noise reduction. In this thesis we focus on single-microphone additive noise reduction and aim at methods that work in the discrete Fourier transform (DFT) domain. The main objective of the presented research is to improve on existing single-microphone schemes for an extended range of noise types and noise levels, thereby making these methods more suitable for mobile speech communication applications than state-of-the-art algorithms. The research topics in this thesis are three-fold. At first, we focus on improved estimation of the a priori signal-to-noise ratio (SNR) from the noisy speech. We focus on two aspects of a priori SNR estimation. Firstly, we present an adaptive time-segmentation algorithm, which we use to reduce the variance of the estimated a priori SNR. Secondly, an approach is presented to reduce the bias of the estimated a priori SNR, which is often present during transitions between speech sounds. Secondly, we investigate the derivation of clean speech estimators under models that take properties of speech into account. This problem is approached from two different angles. At first, we consider the derivation of clean speech estimators under the use of a combined stochastic/deterministic model for the complex DFT coefficients. The use of a deterministic model is based on the fact that certain speech sounds have a more deterministic character. Secondly, we focus on the derivation of complex DFT and magnitude DFT estimators under super-Gaussian densities. Derivation of clean speech estimators under these types of densities is based on measured histograms of speech DFT coefficients. We present two different type of estimators under super-Gaussian densities. Minimum mean-square error (MMSE) estimators are derived under a generalized Gamma density for the clean speech DFT coefficients and DFT magnitudes. Maximum a posteriori (MAP) estimators are derived under the multivariate normal inverse Gaussian (MNIG) density for the clean speech DFT coefficients. Estimators derived under the MNIG density have some theoretical advantages over estimators derived under the generalized Gamma density. More specifically, under the MNIG density the statistical models in the complex DFT and the polar domain are consistent, which is not the case for estimators derived under the generalized Gamma density. In addition, the MNIG density can model vector processes, which allows for taking into account the dependency between the real and imaginary part of DFT coefficients. Finally, we developed a method for tracking of the noise power spectral density (PSD). The developed method is based on the eigenvalue decomposition of correlation matrices that are constructed from time series of noisy DFT coefficients. This approach makes it possible, in contrast to existing methods, to update the noise PSD when speech is continuously present. Furthermore, the tracking delay is considerably reduced compared to state-of-the-art noise tracking algorithms. A comparison is performed between a combination of individual components presented in this thesis and a state-of-the-art speech enhancement system from literature. Subjective experiments by means of a listening test show that the system based on contributions of this thesis improves significantly over the state-of-the-art speech enhancement system.Electrical Engineering, Mathematics and Computer Scienc

    Sound localization in audio-based games for visually impaired children

    No full text
    This thesis describes the design of a sound localization algorithm in audio-based games for visually impaired children. An algorithm was developed to allow real-time audio playback and manipulation, using overlapping audio blocks multiplied by a window function. The audio signal was played through headphones. Multiple sound localization cues are evaluated. Based on these cues, two basic sound localization algorithms are implemented. The first uses only binaural cues. The second expands this with spectral cues by using the head-related impulse response (HRIR). The HRIR was taken from a database, interpolated to obtain optimal resolution and truncated to minimize memory usage. Both localization algorithms work well for lateral localization, but front-back confusions using the HRIR are still common. The signal received from a sound source changes with the distance to the sound source. Both the distance attenuation and propagation delay are implemented. As an alternative means of resolving front-back ambiguities, the use of a head tracker was investigated. Experiments with a webcam based head tracker show that a head tracker is a powerful tool to resolve front- back confusions. However, the latency of webcam based head trackers is too high to be of practical use.Microelectronics & Computer EngineeringElectrical Engineering, Mathematics and Computer Scienc

    Heart Rate Monitoring using Adaptive Noise Cancellation: 2015-2016 Q4

    No full text
    Electrical Engineering, Mathematics and Computer ScienceMicroelectronic

    Verbetering van verstaanbaarheid in teleconferenties door het gebruik van beamformen

    No full text
    De verstaanbaarheid van de aanwezigen bij teleconferenties is minder dan bij gesprekken waarbij beide partijen zich in \'e\'en ruimte bevinden, aangezien stoorsignalen net zo goed worden opgenomen en verder de ruimtelijke informatie van de spraak verloren gaat. Een systeem werd ontworpen dat real-time bepaalt uit welke richting iemand spreekt, geluid uit overige richtingen onderdrukt en het geluid daarna afspeelt alsof het uit de richting kwam waarvandaan gesproken werd om zo het probleem te verhelpen. Acht microfoons nemen het geluid op, waarna dit geluid in het frequentiedomein wordt verwerkt. Eerst wordt er in vijf richtingen gebeamformd, waarna van elk van deze vijf richtingen de energie wordt bepaald. De richting met de hoogste energie wordt dan gekozen uit als de gewenste. Convolutie met een `head-related transfer function' geeft dan geluid dat lijkt alsof het uit de bepaalde richting komt. Dan wordt door middel van de Overlap-Add-methode het signaal teruggetransformeerd naar het tijdsdomein, waarna het wordt afgespeeld via een koptelefoon. De ontworpen beamformer werkt het beste boven frequenties van 2000 Hz. De richtingbepaling werkt voldoende nauwkeurig en snel. Verder kunnen luisteraars herkennen uit welke richting oorspronkelijk gesproken werd. Aanbevolen wordt in een vervolg een frequentieonafhankelijk beamformalgoritme te gebruiken, een richtingbepaling op basis van kruiscorrelatie toe te passen en een manier te ontwerpen die het geluid ruimtelijk weergeeft door luidsprekers.Multimedia Signal ProcessingElectrical Engineering, Mathematics and Computer Scienc

    Sound Reflections in an audio-based game for visually impaired children

    No full text
    In this thesis a model for sound reflections designed which is to be used in an audio-based game for visually impaired people is described. This project is a part of a three-group bachelor graduation project of six peoplein total. The three groups combined create a working game for visually impaired children, meant to be used to train the user’s ears and brain to visualize their surroundings using sound only. Using the mirror source method, a model for sound reflections in concave 3D rooms is obtained. Important parts of this model are the placement of the mirror sources, the calculation of a line of sight and a model for the energy absorption of a wall when sound is reflected off of it. The placement of mirror sources and a model for the wall absorption were successfully implemented. The calculation of a line of sight is only half finished.The implementation of the model in the game adds a sense of the location and the materials of the walls.The execution time of the model is however too low to run the model completely in real time.Electrical EngineeringElectrical Engineering, Mathematics and Computer Scienc

    Noise PSD Insensitive RTF Estimation in a Reverberant and Noisy Environment

    No full text
    Spatial filtering techniques typically rely on estimates of the target relative transfer function (RTF). However, the target speech signal is typically corrupted by late reverberation and ambient noise, which complicates RTF estimation. Existing methods subtract the noise covariance matrix to obtain the target plus late reverberation covariance matrix, from where the RTF is estimated. However, the noise covariance matrix is typically unknown. More specifically, the noise power spectral density (PSD) is typically unknown, while the spatial coherence matrix can be assumed known as it might remain time-invariant for a longer time. Using the spatial coherence matrices we simplify the signal model such that the off-diagonal elements are not affected by the PSDs of the late reverberation and the ambient noise. Then we use these elements to estimate the target covariance matrix, from where the RTF can be obtained. Hence, the resulting estimate of the RTF is insensitive to the noise PSD. Experiments demonstrate the estimation performance of our proposed method.Green Open Access added to TU Delft Institutional Repository 'You share, we take care!' - Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.Signal Processing System
    corecore