Fetching the paper…
Reading the bibliography…
It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Transactions on acoustics, speech, and signal processing , vol. 27, no. 2, pp. 113–120, 1979
1979
Earlier work this paper cites.
J. Allen and D. Berkley, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America , vol. 65, no. 4, pp. 943–950, 1979
1979
Earlier work this paper cites.
R. Lyon, “A computational model of binaural localization and separation,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 8, 1983, pp. 1148–1151
1983
Earlier work this paper cites.
D. B. Paul and J. Baker, “The design for the wall street journal-based CSR corpus,” in Proceedings of the workshop on Speech and Natural Language , 1992, pp. 357–362
1992
Earlier work this paper cites.
O. Cappe, “Elimination of the musical noise phenomenon with the Ephraim and Malah noise suppressor,” IEEE Transactions on Speech and Audio Processing , vol. 2, no. 2, pp. 345–349, 1994
1994
Earlier work this paper cites.
K. Maekawa, H. Koiso, S. Furui, and H. Isahara, “Spontaneous speech corpus of Japanese,” in LREC , vol. 6, 2000, pp. 1–5
2000
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2001, pp. 749–752
2001
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 7, pp. 2125–2136, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann et al. , “The Kaldi speech recognition toolkit,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The diverse environments multi-channel acoustic noise database: A database of multichannel environmental noise recordings,” The Journal of the Acoustical Society of America , vol. 133, no. 5, pp. 3591–3591, 2013
2013
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International Conference on Machine Learning (ICML) , 2013, pp. 1310–1318
2013
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2015, pp. 504–511
2015
Earlier work this paper cites.
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu et al. , “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015, pp. 436–443
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
D. Yu and L. Deng, Automatic speech recognition . Springer, 2016
2016
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Cited alongside, same era.
T. Higuchi, N. Ito, T. Yoshioka, and T. Nakatani, “Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 5210–5214
2016
Cited alongside, same era.
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 196–200
2016
Cited alongside, same era.
Z.-Q. Wang and D. Wang, “A joint training framework for robust automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 4, pp. 796–806, 2016
2016
Cited alongside, same era.
R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix, and T. Nakatani, “Far-field automatic speech recognition,” Proceedings of the IEEE , vol. 109, no. 2, pp. 124–148, 2020
2020
Later among the works it cites.
J. Woo, M. Mimura, K. Yoshii, and T. Kawahara, “End-to-end music-mixed speech recognition,” in Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA) , 2020, pp. 800–804
2020
Later among the works it cites.
T. von Neumann, C. Boeddeker, L. Drude, K. Kinoshita, M. Delcroix, T. Nakatani, and R. Haeb-Umbach, “Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR,” in Interspeech , 2020, pp. 3097–3101
2020
Later among the works it cites.
Q. Wang, I. L. Moreno, M. Saglam, K. Wilson, A. Chiao, R. Liu, Y. He et al. , “Voicefilter-lite: Streaming targeted voice separation for on-device speech recognition,” Interspeech , pp. 2677–2681, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang et al. , “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Interspeech , 2016, pp. 2751–2755
2016
Cited alongside, same era.
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,” Computer Speech & Language , vol. 46, pp. 535–557, 2017
2017
Cited alongside, same era.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Analysis and outcomes,” Computer Speech & Language , vol. 46, pp. 605–626, 2017
2017
Cited alongside, same era.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in Interspeech , 2018, pp. 1561–1565
2018
Cited alongside, same era.
C. Boeddeker, H. Erdogan, T. Yoshioka, and R. Haeb-Umbach, “Exploring practical aspects of neural mask-based beamforming for far-field speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 6697–6701
2018
Cited alongside, same era.
C. Boeddeker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, and R. Haeb-Umbach, “Front-end processing for the CHiME-5 dinner party scenario,” in International Workshop on Speech Processing in Everyday Environments (CHiME-5 Workshop) , vol. 1, 2018
2018
Cited alongside, same era.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Cited alongside, same era.
S.-J. Chen, A. S. Subramanian, H. Xu, and S. Watanabe, “Building state-of-the-art distant speech recognition using the CHiME-4 challenge with a setup of speech enhancement baseline,” in Interspeech , 2018, pp. 1571–1575
2018
Cited alongside, same era.
K. Kinoshita, T. Ochiai, M. Delcroix, and T. Nakatani, “Improving noise robust automatic speech recognition with single-channel time-domain enhancement network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7009–7013
2020
Later among the works it cites.
M. Pariente, S. Cornell, J. Cosentino, S. Sivasankaran, E. Tzinis, J. Heitkaemper, M. Olvera et al. , “Asteroid: the PyTorch-based audio source separation toolkit for researchers,” in Interspeech , 2020, pp. 2637–2641
2020
Later among the works it cites.
H. Sato, T. Ochiai, M. Delcroix, K. Kinoshita, T. Moriya, and N. Kamo, “Should we always separate?: Switching between enhanced and observed signals for overlapping speech recognition,” in Interspeech , 2021, pp. 1149–1153
2021
Later among the works it cites.
C. Boeddeker, W. Zhang, T. Nakatani, K. Kinoshita, T. Ochiai, M. Delcroix, N. Kamo, Y. Qian, and R. Haeb-Umbach, “Convolutive transfer function invariant SDR training criteria for multi-channel reverberant speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 8428–8432
2021
Later among the works it cites.
D. Diaz-Guerra, A. Miguel, and J. R. Beltran, “gpuRIR: A python library for room impulse response simulation with GPU acceleration,” Multimedia Tools and Applications , vol. 80, no. 4, pp. 5653–5671, 2021
2021
Later among the works it cites.
J. Li, “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing , vol. 11, no. 1, 2022
2022
Later among the works it cites.
X. Chang, T. Maekaku, Y. Fujita, and S. Watanabe, “End-to-end integration of speech recognition, speech enhancement, and self-supervised learning representation,” in Interspeech , 2022, pp. 3819–3823
2022
Later among the works it cites.
K. Iwamoto, T. Ochiai, M. Delcroix, R. Ikeshita, H. Sato, S. Araki, and S. Katagiri, “How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR,” in Interspeech , 2022, pp. 5418–5422
2022
Later among the works it cites.
C. Zorila and R. Doddipatla, “Speaker reinforcement using target source extraction for robust automatic speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6297–6301
2022
Later among the works it cites.
H. Sato, T. Ochiai, M. Delcroix, K. Kinoshita, N. Kamo, and T. Moriya, “Learning to enhance or not: Neural network-based switching of enhanced and observed signals for overlapping speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6287–6291
2022
Later among the works it cites.
Y. Koizumi, S. Karita, A. Narayanan, S. Panchapagesan, and M. Bacchiani, “SNRi target training for joint speech enhancement and recognition,” in Interspeech , 2022, pp. 1173–1177
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, pp. 1505–1518, 2022
2022
Later among the works it cites.
K. Zmolikova, M. Delcroix, T. Ochiai, K. Kinoshita, J. Černockỳ, and D. Yu, “Neural target speech extraction: An overview,” IEEE Signal Processing Magazine , vol. 40, no. 3, pp. 8–29, 2023
2023
Later among the works it cites.