SDJul 2

Speaker head orientation estimation with a single microphone array using phase spectrogram features

arXiv:2607.021292.8
Predicted impact top 86% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for robust head orientation estimation from audio in smart environments and driver monitoring, offering a novel method that outperforms prior approaches.

The paper proposes a deep neural network using phase spectrogram features from a single microphone array to estimate speaker head orientation, achieving state-of-the-art accuracy with a mean angular error of 11.3 degrees after personalization.

Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverages the phase component of the short-time Fourier transform from a single microphone array as input to a deep neural network combining convolutional, recurrent, and self-attention layers. Unlike prior methods that use physics-informed handcrafted features or raw waveform inputs, our approach enables robust learning from simulated and real data. Trained on a large-scale dataset generated with voice directivity patterns and fine-tuned on real recordings, our model achieves state-of-the-art accuracy, outperforming baselines under both clean and noisy conditions. Personalization experiments further demonstrate significant gains, reaching a mean angular error of 11.3 degrees when adapting to individual users and environments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes