Fetching the paper…
Reading the bibliography…
Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs.
“Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,”
S. Davis and P. Mermelstein, · 1980
Earlier work this paper cites.
“Csr-i (wsj0) complete ldc93s6a.,”
et al. Garofolo, John S., · 1993
Earlier work this paper cites.
“Comments on” noise injection into inputs in back propagation learning”,”
Yves Grandvalet and Stéphane Canu, · 1995
Earlier work this paper cites.
“Training with noise is equivalent to tikhonov regularization,”
Chris M Bishop, · 1995
Earlier work this paper cites.
“The effects of adding noise during backpropagation training on a generalization performance,”
Guozhong An, · 1996
Earlier work this paper cites.
“A compact model for speaker-adaptive training,”
Tasos Anastasakos, John McDonough, Richard Schwartz, and John Makhoul, · 1996
Earlier work this paper cites.
“Noise injection: Theoretical prospects,”
Yves Grandvalet, Stéphane Canu, and Stéphane Boucheron, · 1997
Earlier work this paper cites.
“A recursive feature vector normalization approach for robust speech recognition in noise,”
Olli Viikki, David Bye, and Kari Laurila, · 1998
Earlier work this paper cites.
“Quicknet,”
ICSI Berkeley, · 2000
Earlier work this paper cites.
“Adaptive training with joint uncertainty decoding for robust recognition of noisy data,”
Hank Liao and MJ F Gales, · 2007
Earlier work this paper cites.
“Multi-condition training for unknown environment adaptation in robust asr under real conditions,”
J Rajnoha, · 2009
Earlier work this paper cites.
“Exemplar-based sparse representations for noise robust automatic speech recognition,”
Jort F Gemmeke, Tuomas Virtanen, and Antti Hurmalainen, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
“Recurrent neural networks for noise reduction in robust asr,”
Andrew L Maas, Quoc V Le, Tyler M O’Neil, Oriol Vinyals, Patrick Nguyen, and Andrew Y Ng, · 2012
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“An investigation of deep neural networks for noise robust speech recognition,”
Michael L Seltzer, Dong Yu, and Yongqiang Wang, · 2013
Cited alongside, same era.
“Audio-visual deep learning for noise robust speech recognition,”
Jing Huang and Brian Kingsbury, · 2013
Cited alongside, same era.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Later among the works it cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Later among the works it cites.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Later among the works it cites.
“A statistical analysis on the impact of noise on mfcc features for speech recognition,”
U. Bhattacharjee, Swapnanil Gogoi, and Rubi Sharma, · 2016
Later among the works it cites.
“Invariant representations for noisy speech recognition,”
Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel, Bhuvana Ramabhadran, Samuel Thomas, and Yoshua Bengio, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Cited alongside, same era.
“An overview of noise-robust automatic speech recognition,”
Jinyu Li, Li Deng, Yifan Gong, and Reinhold Haeb-Umbach, · 2014
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, and et al., · 2015
Cited alongside, same era.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Cited alongside, same era.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D Manning, · 2015
Cited alongside, same era.
“Musan: A music, speech, and noise corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Cited alongside, same era.
“Domain-adversarial training of neural networks,”
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky, · 2016
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Trace norm regularization and faster inference for embedded speech recognition rnns,”
Markus Kliegl, Siddharth Goyal, Kexin Zhao, Kavya Srinet, and Mohammad Shoeybi, · 2017
Later among the works it cites.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Closest in time.
“Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin chinese,”
Shiyou Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Closest in time.
“Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Closest in time.
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Closest in time.
“Adversarial logit pairing,”
H. Kannan, A. Kurakin, and I. Goodfellow, · 2018
Closest in time.