Fetching the paper…
Reading the bibliography…
We investigate the effect of speaker localization on the performance of speech recognition systems in a multispeaker, multichannel environment.
“The generalized correlation method for estimation of time delay,”
C. Knapp and G. Carter, · 1976
Earlier work this paper cites.
“Computational auditory scene analysis,”
G.J. Brown and Martin Cooke, · 1994
Earlier work this paper cites.
Microphone Arrays: Signal Processing Techniques and Applications
Michael Brandstein and Darren Ward, Eds., · 2001
Earlier work this paper cites.
“Spatially pre-processed speech distortion weighted multi-channel Wiener filtering for noise reduction,”
A. Spriet, M. Moonen, and J. Wouters, · 2004
Earlier work this paper cites.
Computational Auditory Scene Analysis: Principles, Algorithms, and Applications
DeLiang Wang and Guy J. Brown, · 2006
Earlier work this paper cites.
“Blind acoustic beamforming based on generalized eigenvalue decomposition,”
E. Warsitz and R. Haeb-Umbach, · 2007
Earlier work this paper cites.
Distant Speech Recognition
M. Wölfel and J. McDonough, · 2009
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“On the impact of localization errors on HRTF-based robust least-squares beamforming,”
Hendrik Barfuss and Walter Kellermann, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Deep attractor network for single-microphone speaker separation,”
Zhuo Chen, Yi Luo, and Nima Mesgarani, · 2017
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Cited alongside, same era.
“A consolidated perspective on multimicrophone speech enhancement and source separation,”
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, · 2017
Cited alongside, same era.
“DOA-informed source extraction in the presence of competing talkers and background noise,”
Maja Taseska and Emanuël A. P. Habets, · 2017
Cited alongside, same era.
“The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,”
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal, · 2018
Later among the works it cites.
“Keyword-based speaker localization: Localizing a target speaker in a multi-speaker environment,”
Sunit Sivasankaran, Emmanuel Vincent, and Dominique Fohr, · 2018
Later among the works it cites.
“Rank-1 constrained multichannel Wiener filter for speech recognition in noisy environments,”
Ziteng Wang, Emmanuel Vincent, Romain Serizel, and Yonghong Yan, · 2018
Later among the works it cites.
“RIR-Generator: Room impulse response generator,”
Emanuël A. P. Habets, · 2018
Later among the works it cites.
“Multi-Channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,”
Zhong-Qiu Wang, Jonathan Le Roux, and John R. Hershey, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Cracking the cocktail party problem by multi-beam deep attractor network,”
Z. Chen, J. Li, X. Xiao, T. Yoshioka, H. Wang, Z. Wang, and Y. Gong, · 2017
Cited alongside, same era.
“Listening to each speaker one by one with recurrent selective hearing networks,”
Keisuke Kinoshita, Lukas Drude, Marc Delcroix, and Tomohiro Nakatani, · 2018
Cited alongside, same era.
“Multichannel speech separation with recurrent neural networks from high-order ambisonics recordings,”
L. Perotin, R. Serizel, E. Vincent, and A. Guérin, · 2018
Cited alongside, same era.
“Multi-Channel Overlapped Speech Recognition with Location Guided Speech Extraction Network,”
Z. Chen, X. Xiao, T. Yoshioka, H. Erdogan, J. Li, and Y. Gong, · 2018
Cited alongside, same era.
“Combining spectral and spatial features for deep learning based blind speaker separation,”
Z. Wang and D. Wang, · 2019
Closest in time.
“Analysis of deep clustering as preprocessing for automatic speech recognition of sparsely overlapping speech,”
Tobias Menne, Ilya Sklyar, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“The second DIHARD diarization challenge: Dataset, task, and baselines,”
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia, Jun Du, Sriram Ganapathy, and Mark Liberman, · 2019
Closest in time.