Fetching the paper…
Reading the bibliography…
It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with single-channel speech enhancement (SE).
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Transactions on acoustics, speech, and signal processing , vol. 27, no. 2, pp. 113–120, 1979
1979
Earlier work this paper cites.
J. Allen and D. Berkley, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America , vol. 65, no. 4, pp. 943–950, 1979
1979
Earlier work this paper cites.
R. Lyon, “A computational model of binaural localization and separation,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 8, 1983, pp. 1148–1151
1983
Earlier work this paper cites.
H. Erdogan, J. Hershey, S. Watanabe, M. Mandel, and J. Le Roux, “Improved MVDR beamforming using single-channel mask prediction networks,” in Interspeech , 2016, pp. 1981–1985
1985
Earlier work this paper cites.
O. Cappe, “Elimination of the musical noise phenomenon with the ephraim and malah noise suppressor,” IEEE transactions on Speech and Audio Processing , vol. 2, no. 2, pp. 345–349, 1994
1994
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
J. Garofalo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete,” Linguistic Data Consortium, Philadelphia , 2007
2007
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann et al. , “The Kaldi speech recognition toolkit,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International Conference on Machine Learning (ICML) , 2013, pp. 1310–1318
2013
Cited alongside, same era.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2015, pp. 504–511
2015
Cited alongside, same era.
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu et al. , “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015, pp. 436–443
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in Interspeech , 2018, pp. 1561–1565
2018
Later among the works it cites.
S.-J. Chen, A. S. Subramanian, H. Xu, and S. Watanabe, “Building state-of-the-art distant speech recognition using the chime-4 challenge with a setup of speech enhancement baseline,” in Interspeech , 2018, pp. 1571–1575
2018
Later among the works it cites.
T. Menne, R. Schlüter, and H. Ney, “Investigation into joint optimization of single channel speech enhancement and acoustic modeling for robust asr,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6660–6664
2019
Later among the works it cites.
M. Fujimoto and H. Kawai, “One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features.” in Interspeech , 2019, pp. 486–490
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 196–200
2016
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Interspeech , 2016, pp. 2751–2755
2016
Cited alongside, same era.
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,” Computer Speech & Language , vol. 46, pp. 535–557, 2017
2017
Cited alongside, same era.
C. Boeddeker, H. Erdogan, T. Yoshioka, and R. Haeb-Umbach, “Exploring practical aspects of neural mask-based beamforming for far-field speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 6697–6701
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Later among the works it cites.
J. L. Roux, S. Wisdom, H. Erdogan, and J. Hershey, “SDR – half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 626–630
2019
Later among the works it cites.
K. Kinoshita, T. Ochiai, M. Delcroix, and T. Nakatani, “Improving noise robust automatic speech recognition with single-channel time-domain enhancement network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7009–7013
2020
Later among the works it cites.
M. Pariente, S. Cornell, J. Cosentino, S. Sivasankaran, E. Tzinis, J. Heitkaemper, M. Olvera et al. , “Asteroid: the PyTorch-based audio source separation toolkit for researchers,” in Interspeech , 2020, pp. 2637–2641
2020
Later among the works it cites.