Fetching the paper…
Reading the bibliography…
Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech distortion.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 2, pp. 113–120, 1979
1979
Earlier work this paper cites.
A. Varga and H. J. Steeneken, “Assessment for automatic speech recognition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” ELSEVIER Speech Commun. , vol. 12, no. 3, pp. 247–251, 1993
1993
Earlier work this paper cites.
P. Scalart and J. Filho, “Speech enhancement based on a priori signal to noise estimation,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , vol. 2, 1996, pp. 629–632
1996
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Int. Conf. Machine Learning (ICML) , 2006, pp. 369–376
2006
Earlier work this paper cites.
H. Jegou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” IEEE Trans. Pattern Analysis & Machine Intelligence , vol. 33, no. 1, pp. 117–128, 2010
2010
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag. , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proc. 21st ACM Int. Conf. Multimedia , 2013, pp. 411–412
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in Int. Conf. Machine Learning (ICML) , 2014, pp. 1764–1772
2014
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Proc. of NIPS , vol. 28, 2015, pp. 577–585
2015
Earlier work this paper cites.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. L. Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,” in Proc. of LVA/ICA . Springer, 2015, pp. 91–99
2015
Earlier work this paper cites.
K. Han, Y. Wang, D. Wang, W. S. Woods, I. Merks, and T. Zhang, “Learning spectral mapping for speech dereverberation and denoising,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 23, no. 6, pp. 982–992, 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
Z.-Q. Wang and D. Wang, “A joint training framework for robust automatic speech recognition,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 24, no. 4, pp. 796–806, 2016
2016
Earlier work this paper cites.
T. Menne, J. Heymann, A. Alexandridis, K. Irie, A. Zeyer, M. Kitza, P. Golik, I. Kulikov, L. Drude, R. Schlüter, and et al., “The rwth/upb/forth system combination for the 4th chime challenge evaluation,” in Proc. of CHiME-4 Workshop , 2016, pp. 49–51
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in Int. Conf. Machine Learning (ICML) , 2016, pp. 173–182
2016
Earlier work this paper cites.
J. Du, Y.-H. Tu, L. Sun, F. Ma, H.-K. Wang, J. Pan, C. Liu, J.-D. Chen, and C.-H. Lee, “The ustc-iflytek system for chime-4 challenge,” Proc. CHiME , vol. 4, pp. 36–38, 2016
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 4835–4839
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. of NIPS , vol. 30, pp. 6000–6010, 2017
2017
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in Proc. of ICLR , 2017
2017
Earlier work this paper cites.
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, “The microsoft 2017 conversational speech recognition system,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5934–5938
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5884–5888
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Tasnet: Time-domain audio separation network for real-time, single-channel speech separation,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 696–700
2018
Cited alongside, same era.
M. H. Soni, N. Shah, and H. A. Patil, “Time-frequency masking-based speech enhancement using generative adversarial network,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5039–5043
2018
Cited alongside, same era.
S. Ling, Y. Liu, J. Salazar, and K. Kirchhoff, “Deep contextualized acoustic representations for semi-supervised speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 6429–6433
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. of NIPS , 2020, pp. 12 449–12 460
2020
Later among the works it cites.
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord, “Learning robust and multilingual speech representations,” in Empirical Methods in Natural Language Process.: Findings , 2020, pp. 1182–1192
2020
Later among the works it cites.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 7414–7418
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, “Visualizing the loss landscape of neural nets,” Proc. of NIPS , vol. 31, 2018
2018
Cited alongside, same era.
A. Pandey and D. Wang, “Tcnn: Temporal convolutional neural network for real-time speech enhancement in the time domain,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2019, pp. 6875–6879
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
A. Pandey and D. Wang, “A new framework for cnn-based speech enhancement in the time domain,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 7, pp. 1179–1188, 2019
2019
Cited alongside, same era.
M. Fujimoto and H. Kawai, “One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features.” in ISCA Interspeech , 2019, pp. 486–490
2019
Cited alongside, same era.
B. Liu, S. Nie, S. Liang, W. Liu, M. Yu, L. Chen, S. Peng, C. Li et al. , “Jointly adversarial enhancement training for robust end-to-end speech recognition.” in ISCA Interspeech , 2019, pp. 491–495
2019
Cited alongside, same era.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in ISCA Interspeech , 2019, pp. 146–150
2019
Cited alongside, same era.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 6989–6993
2020
Later among the works it cites.
A. T. Liu, S.-W. Li, and H.-y. Lee, “Tera: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 2351–2366, 2021
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Prasad, P. Jyothi, and R. Velmurugan, “An investigation of end-to-end models for robust speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 6893–6897
2021
Later among the works it cites.
Y. Hu, N. Hou, C. Chen, and E. Siong Chng, “Interactive feature fusion for end-to-end noise-robust speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 6292–6296
2022
Closest in time.
2022
Closest in time.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li, and X. Yu, “Unispeech-sat: Universal speech representation learning with speaker aware pre-training,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 6152–6156
2022
Closest in time.
Y. Wang, J. Li, H. Wang, Y. Qian, C. Wang, and Y. Wu, “Wav2vec-switch: Contrastive learning from original-noisy speech pairs for robust speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 7097–7101
2022
Closest in time.
H. Wang, Y. Qian, X. Wang, Y. Wang, C. Wang, S. Liu, T. Yoshioka, J. Li, and D. Wang, “Improving noise robustness of contrastive speech representation learning with speech reconstruction,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 6062–6066
2022
Closest in time.
Q.-S. Zhu, J. Zhang, Z.-Q. Zhang, M.-H. Wu, X. Fang, and L.-R. Dai, “A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 3174–3178
2022
Closest in time.