Fetching the paper…
Reading the bibliography…
Recently, the end-to-end approach has proven its efficacy in monaural multi-speaker speech recognition.
E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the Acoustical Society of America , vol. 25, no. 5, 1953
1953
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 7, Sep. 2007
2007
Earlier work this paper cites.
M. Souden, J. Benesty, and S. Affes, “On optimal frequency-domain multichannel linear filtering for noise reduction,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 2, 2009
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proc. International Conference on Machine Learning (ICML) , Jun. 2009
2009
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in Proc. International Conference on Machine Learning (ICML) , Jun. 2014
2014
Earlier work this paper cites.
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu, W. J. Fabian, M. Espi, T. Higuchi et al. , “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices,” in Proc. IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , Dec. 2015
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2016
2016
Earlier work this paper cites.
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in Proc. ISCA Interspeech , Sep. 2016
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2016
2016
Earlier work this paper cites.
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2016
2016
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, and J. Le Roux, “Improved MVDR beamforming using single-channel mask prediction networks,” in Proc. ISCA Interspeech , Sep. 2016
2016
Earlier work this paper cites.
T. Menne, J. Heymann, A. Alexandridis, K. Irie, A. Zeyer, M. Kitza, P. Golik, I. Kulikov, L. Drude, R. Schlüter et al. , “The RWTH/UPB/FORTH system combination for the 4th CHiME challenge evaluation,” in Proc. CHiME workshop , Sep. 2016
2016
Earlier work this paper cites.
D. Amodei et al. , “Deep Speech 2: End-to-end speech recognition in English and Mandarin,” in Proc. International Conference on Machine Learning (ICML) , Jun. 2016
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2017
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 25, no. 10, 2017
2017
Cited alongside, same era.
D. Yu, X. Chang, and Y. Qian, “Recognizing multi-talker speech with permutation invariant training,” in Proc. ISCA Interspeech , Aug. 2017
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2017
2017
Cited alongside, same era.
T. Yoshioka, H. Erdogan, Z. Chen, X. Xiao, and F. Alleva, “Recognizing overlapped speech in meetings: A multichannel separation approach using neural networks,” in Proc. ISCA Interspeech , Sep. 2018
2018
Later among the works it cites.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018
2018
Later among the works it cites.
S. Braun, D. Neil, J. Anumula, E. Ceolini, and S.-C. Liu, “Multi-channel attention for end-to-end speech recognition.” in Proc. ISCA Interspeech , Sep. 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Hori, S. Watanabe, and J. Hershey, “Joint CTC/attention decoding for end-to-end speech recognition,” in Proc. Annual Meeting of the Association for Computational Linguistics (ACL) , vol. 1, Jul. 2017
2017
Cited alongside, same era.
J. Heymann, L. Drude, C. Boeddeker, P. Hanebrink, and R. Haeb-Umbach, “Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2017
2017
Cited alongside, same era.
T. Ochiai, S. Watanabe, T. Hori, and J. R. Hershey, “Multichannel end-to-end speech recognition,” in Proc. International Conference on Machine Learning (ICML) , 2017
2017
Cited alongside, same era.
T. Ochiai, S. Watanabe, and S. Katagiri, “Does speech enhancement work with end-to-end ASR objectives?: Experimental analysis of multichannel end-to-end ASR,” in Proc. International Workshop on Machine Learning for Signal Processing (MLSP) , Sep. 2017
2017
Cited alongside, same era.
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimicrophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 25, no. 4, 2017
2017
Cited alongside, same era.
S. Settle, J. Le Roux, T. Hori, S. Watanabe, and J. R. Hershey, “End-to-end multi-speaker speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018
2018
Cited alongside, same era.
Y. Qian, X. Chang, and D. Yu, “Single-channel multi-talker speech recognition with permutation invariant training,” Speech Communication , vol. 104, 2018
2018
Cited alongside, same era.
H. Seki, T. Hori, S. Watanabe, J. Le Roux, and J. R. Hershey, “A purely end-to-end system for multi-speaker speech recognition,” in Proc. Annual Meeting of the Association for Computational Linguistics (ACL) , Jul. 2018
2018
Cited alongside, same era.
T. Hori, J. Cho, and S. Watanabe, “End-to-end speech recognition with word-based RNN language models,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , Dec. 2018
2018
Later among the works it cites.
L. Drude, J. Heymann, C. Boeddeker, and R. Haeb-Umbach, “NARA-WPE: A Python package for weighted prediction error dereverberation in Numpy and Tensorflow for online and offline processing,” in ITG Fachtagung Sprachkommunikation (ITG) , Oct. 2018
2018
Later among the works it cites.
2019
Closest in time.
X. Chang, Y. Qian, K. Yu, and S. Watanabe, “End-to-end monaural multi-speaker ASR system without pretraining,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2019
2019
Closest in time.
W. Minhua, K. Kumatani, S. Sundaram, N. Ström, and B. Hoffmeister, “Frequency domain multi-channel acoustic modeling for distant speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2019
2019
Closest in time.
X. Wang, R. Li, S. H. Mallidi, T. Hori, S. Watanabe, and H. Hermansky, “Stream attention-based multi-array end-to-end speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2019
2019
Closest in time.
2019
Closest in time.
J. Le Roux, S. T. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – half-baked or well done?” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2019
2019
Closest in time.