Joerg Schmalenstroeer

AS
h-index2
3papers
32citations
Novelty48%
AI Score29

3 Papers

3.3ASOct 25, 2021Code
On Synchronization of Wireless Acoustic Sensor Networks in the Presence of Time-varying Sampling Rate Offsets and Speaker Changes

Tobias Gburrek, Joerg Schmalenstroeer, Reinhold Haeb-Umbach

A wireless acoustic sensor network records audio signals with sampling time and sampling rate offsets between the audio streams, if the analog-digital converters (ADCs) of the network devices are not synchronized. Here, we introduce a new sampling rate offset model to simulate time-varying sampling frequencies caused, for example, by temperature changes of ADC crystal oscillators, and propose an estimation algorithm to handle this dynamic aspect in combination with changing acoustic source positions. Furthermore, we show how deep neural network based estimates of the distances between microphones and human speakers can be used to determine the sampling time offsets. This enables a synchronization of the audio streams to reflect the physical time differences of flight.

1.2ASDec 11, 2020Code
Iterative Geometry Calibration from Distance Estimates for Wireless Acoustic Sensor Networks

Tobias Gburrek, Joerg Schmalenstroeer, Reinhold Haeb-Umbach

In this paper we present an approach to geometry calibration in wireless acoustic sensor networks, whose nodes are assumed to be equipped with a compact microphone array. The proposed approach solely works with estimates of the distances between acoustic sources and the nodes that record these sources. It consists of an iterative weighted least squares localization procedure, which is initialized by multidimensional scaling. Alongside the sensor node locations, also the positions of the acoustic sources are estimated. Furthermore, we derive the Cramer-Rao lower bound (CRLB) for source and sensor position estimation, and show by simulation that the estimator is efficient.

4.3ASMay 20, 2020
Statistical and Neural Network Based Speech Activity Detection in Non-Stationary Acoustic Environments

Jens Heitkaemper, Joerg Schmalenstroeer, Reinhold Haeb-Umbach

Speech activity detection (SAD), which often rests on the fact that the noise is "more" stationary than speech, is particularly challenging in non-stationary environments, because the time variance of the acoustic scene makes it difficult to discriminate speech from noise. We propose two approaches to SAD, where one is based on statistical signal processing, while the other utilizes neural networks. The former employes sophisticated signal processing to track the noise and speech energies and is meant to support the case for a resource efficient, unsupervised signal processing approach. The latter introduces a recurrent network layer that operates on short segments of the input speech to do temporal smoothing in the presence of non-stationary noise. The systems are tested on the Fearless Steps challenge, which consists of the transmission data from the Apollo-11 space mission. The statistical SAD achieves comparable detection performance to earlier proposed neural network based SADs, while the neural network based approach leads to a decision cost function of 1.07% on the evaluation set of the 2020 Fearless Steps Challenge, which sets a new state of the art.