Fetching the paper…
Reading the bibliography…
Sequence-to-sequence (S2S) modeling is becoming a popular paradigm for automatic speech recognition (ASR) because of its ability to jointly optimize all the conventional ASR components in an end-to-end (E2E) fashion.
B. Li, T. N. Sainath, R. J. Weiss, K. W. Wilson, and M. Bacchiani, “Neural network adaptive beamforming for robust multichannel speech recognition,” in
1980
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, and J. Le Roux, “Improved MVDR beamforming using single-channel mask prediction networks.” in
1985
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based CSR corpus,” in
1992
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,”
2007
Earlier work this paper cites.
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, and B. Juang, “Speech dereverberation based on variance-normalized delayed linear prediction,”
2010
Earlier work this paper cites.
M. Souden, J. Benesty, and S. Affes, “On optimal frequency-domain multichannel linear filtering for noise reduction,”
2010
Earlier work this paper cites.
G. Hinton
2012
Earlier work this paper cites.
T. Yoshioka and T. Nakatani, “Generalization of multi-channel linear prediction methods for blind mimo impulse response shortening,”
2012
Earlier work this paper cites.
J. Li, L. Deng, Y. Gong, and R. Haeb-Umbach, “An overview of noise-robust automatic speech recognition,”
2014
Earlier work this paper cites.
Y. Tachioka, T. Narita, F. J. Weninger, and S. Watanabe, “Dual system combination approach for various reverberant environments with dereverberation techniques,” in
2014
Earlier work this paper cites.
M. J. Alam, V. Gupta, P. Kenny, and P. Dumouchel, “Use of multiple front-ends and i-vector-based speaker adaptation for robust speech recognition,”
2014
Earlier work this paper cites.
M. Ravanelli
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Cited alongside, same era.
D. Amodei
2016
Cited alongside, same era.
K. Kinoshita
2016
Cited alongside, same era.
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in
2016
Cited alongside, same era.
T. Menne
2016
Cited alongside, same era.
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in
2017
Later among the works it cites.
C.-C. Chiu
2018
Later among the works it cites.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth CHiME speech separation and recognition challenge: Dataset, task and baselines,” in
2018
Later among the works it cites.
L. Drude
2018
Later among the works it cites.
S. Braun, D. Neil, J. Anumula, E. Ceolini, and S.-C. Liu, “Multi-channel attention for end-to-end speech recognition,” in
2018
Later among the works it cites.
S.-J. Chen, A. S. Subramanian, H. Xu, and S. Watanabe, “Building state-of-the-art distant speech recognition using the CHiME-4 challenge with a setup of speech enhancement baseline,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Ochiai, S. Watanabe, T. Hori, and J. R. Hershey, “Multichannel end-to-end speech recognition,” in
2017
Cited alongside, same era.
T. Ochiai, S. Watanabe, T. Hori, J. R. Hershey, and X. Xiao, “Unified architecture for multichannel end-to-end speech recognition with neural beamforming,”
2017
Cited alongside, same era.
J. Heymann, L. Drude, C. Boeddeker, P. Hanebrink, and R. Haeb-Umbach, “Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system,” in
2017
Cited alongside, same era.
K. Kinoshita, M. Delcroix, H. Kwon, T. Mori, and T. Nakatani, “Neural network-based spectrum estimation for online WPE dereverbertion,” in
2017
Cited alongside, same era.
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimicrophone speech enhancement and source separation,”
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,”
2017
Cited alongside, same era.
2018
Later among the works it cites.
T. Menne, R. Schlüter, and H. Ney, “Speaker adapted beamforming for multi-channel automatic speech recognition,” in
2018
Later among the works it cites.
J. Heymann, L. Drude, R. Haeb-Umbach, K. Kinoshita, and T. Nakatani, “Frame-online DNN-WPE dereverberation,” in
2018
Later among the works it cites.
S. Watanabe
2018
Later among the works it cites.
T. Hori, J. Cho, and S. Watanabe, “End-to-end speech recognition with word-based RNN language models,” in
2018
Later among the works it cites.
L. Drude, J. Heymann, C. Boeddeker, and R. Haeb-Umbach, “NARA-WPE: A Python package for weighted prediction error dereverberation in Numpy and Tensorflow for online and offline processing,” in
2018
Later among the works it cites.
W. Minhua, K. Kumatani, S. Sundaram, N. Strom, and B. Hoffmeister, “Frequency domain multi-channel acoustic modeling for distant speech recognition,” in
2019
Closest in time.