Fetching the paper…
Reading the bibliography…
Despite the strong modeling power of neural network acoustic models, speech enhancement has been shown to deliver additional word error rate improvements if multi-channel data is available.
“Domain adaptation for robust automatic speech recognition in car environments,”
R.-D. Bippus, A. Fischer, and V. Stahl, · 1946
Earlier work this paper cites.
“Neural network adaptive beamforming for robust multichannel speech recognition,”
B. Li, T. N. Sainath, R. J. Weiss, K. W. Wilson, and M. Bacchiani, · 1980
Earlier work this paper cites.
“Improved MVDR beamforming using single-channel mask prediction networks.,”
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, and J. Le Roux, · 1985
Earlier work this paper cites.
“Large-vocabulary speech recognition under adverse acoustic environments,”
L. Deng, A. Acero, M. Plumpe, and X. Huang, · 2000
Earlier work this paper cites.
“A robust and precise method for solving the permutation problem of frequency-domain blind source separation,”
H. Sawada, R. Mukai, S. Araki, and S. Makino, · 2004
Earlier work this paper cites.
“Blind acoustic beamforming based on generalized eigenvalue decomposition,”
E. Warsitz and R. Haeb-Umbach, · 2007
Earlier work this paper cites.
“Acoustic beamforming for speaker diarization of meetings,”
X. Anguera, C. Wooters, and J. Hernando, · 2007
Earlier work this paper cites.
“On optimal frequency-domain multichannel linear filtering for noise reduction,”
M. Souden, J. Benesty, and S. Affes, · 2010
Earlier work this paper cites.
“Recurrent neural network based language model,”
T. Mikolov, M. Karafiát, L. Burget, J. Černocký, and S. Khudanpur, · 2010
Earlier work this paper cites.
“Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening,”
T. Yoshioka and T. Nakatani, · 2012
Cited alongside, same era.
“Is speech enhancement pre-processing still relevant when using deep neural networks for acoustic modeling?,”
M. Delcroix, Y. Kubo, T. Nakatani, and A. Nakamura, · 2013
Cited alongside, same era.
“Sequence-discriminative training of deep neural networks,”
A. Ghoshal and D. Povey, · 2013
Cited alongside, same era.
“Linear prediction-based dereverberation with advanced speech enhancement and recognition technologies for the REVERBED challenge,”
M. Delcroix, T. Yoshioka, A. Ogawa, Y. Kubo, M. Fujimoto, N. Ito, K. Kinoshita, M. Espi, T. Hori, T. Nakatani, and A. Nakamura, · 2014
Cited alongside, same era.
“Convolutional neural networks for speech recognition,”
O. Abdel-Hamid, A.-R. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, · 2014
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, D. Povey, M. Seltzer, and S. Khudanpur, · 2017
Later among the works it cites.
“The fifth CHiME speech separation and recognition challenge: Dataset, task and baselines,”
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, · 2018
Later among the works it cites.
“Front-end processing for the CHiME-5 dinner party scenario,”
C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, and R. Haeb-Umbach, · 2018
Later among the works it cites.
“NARA-WPE: A Python package for weighted prediction error dereverberation in Numpy and Tensorflow for online and offline processing,”
L. Drude, J. Heymann, C. Boeddeker, and R. Haeb-Umbach, · 2018
Later among the works it cites.
“Semi-orthogonal low-rank matrix factorization for deep neural networks,”
D. Povey, G. Cheng, Y. Wang, K. Lia, H. Xu, M. Yarmohammadi, and S. Khudanpur, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The third ‘CHiME’ speech separation and recognition challenge: dataset, task and baselines,”
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, · 2015
Cited alongside, same era.
“An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
E. Vincent, S. Watanabe, A.-A. Nugraha, JonBarker, and RicardMarxer, · 2016
Cited alongside, same era.
“The RWTH /UPB/FORTH system combination for the 4th CHiME challenge evaluation,”
T. Menne, J. Heymann, A. Alexandridis, K. Irie, A. Zeyer, M. Kitza, P. Golik, I. Kulikov, L. Drude, R. Schlüter, H. Ney, R. Haeb-Umbach, and A. Mouchtaris, · 2016
Cited alongside, same era.
“Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array signal processing,”
N. Ito, S. Araki, and T. Nakatani, · 2016
Cited alongside, same era.
“Acoustic modeling for overlapping speech recognition: JHU Chime-5 Challenge system,”
V. Manohar, S.-J. Chen, Z. Wang, Y. Fujita, S. Watanabe, and S. Khudanpur, · 2019
Closest in time.
“On reducing the effect of speaker overlap for CHiME-5,”
C. Zorilă and R. Doddipatla, · 2019
Closest in time.
“Guided source separation meets a strong ASR backend: Hitachi/Paderborn University joint investigation for dinner party ASR,”
N. Kanda, C. Boeddeker, J. Heitkaemper, Y. Fujita, S. Horiguchi, K. Nagamatsu, and R. Haeb-Umbach, · 2019
Closest in time.