Fetching the paper…
Reading the bibliography…
Self-attention has been a huge success for many downstream tasks in NLP, which led to exploration of applying self-attention to speech problems as well.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan R. Salakhutdinov, · 1901
Earlier work this paper cites.
“Sequence-to-sequence speech recognition with time-depth separable convolutions,”
Awni Hannun, Ann Lee, Qiantong Xu, and Ronan Collobert, · 1904
Earlier work this paper cites.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiua, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 1904
Earlier work this paper cites.
“RWTH ASR systems for LibriSpeech: Hybrid vs attention - w/o data augmentation,”
Christoph Luscher, Eugen Beck, Kazuki Irie1, Markus Kitza1, Wilfried Michel, Albert Zeyer, Ralf Schluter, and Hermann Ney, · 1905
Earlier work this paper cites.
“XLNet: Generalized autoregressive pretraining for language understanding,”
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le, · 1906
Earlier work this paper cites.
“A comprative study on Transformer vs RNN in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang, · 1909
Earlier work this paper cites.
“Phoneme recognition using time-delay neural networks,”
Alexander Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J. Lang, · 1989
Earlier work this paper cites.
“Improved backing-off for M-gram language modeling,”
Reinhard Kneser and Hermann Ney, · 1995
Earlier work this paper cites.
“An empirical study of smoothing techniques for language modeling,”
Stanley F. Chen and Joshua Goodman, · 1996
Earlier work this paper cites.
“Maximum likelihood linear transformations for HMM-based speech recognition,”
Mark J. F. Gales, · 1997
Earlier work this paper cites.
“SRILM – An extensible language modeling toolkit,”
A. Stolcke, · 2002
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernandez, Faustino Gomez, and Jurgen Schmidhuber, · 2006
Earlier work this paper cites.
“Joint-sequence models for grapheme-to-phoneme conversion,”
Maximilian Bisani and Hermann Ney, · 2008
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Cited alongside, same era.
“Improving neural networks by preventing co-adaptation of feature detectors,”
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2012
Cited alongside, same era.
“Restructuring of deep neural network acoustic models with singular value decomposition,”
Jian Xue, Jinyu Li, and Yifan Gong, · 2013
Cited alongside, same era.
“On rectified linear units for speech processing,”
Matthew D. Zeiler, Marc’Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Kangye Yang, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, and Geoffrey E. Hinton, · 2013
Cited alongside, same era.
“LibriVox: Free public domain audiobooks,”
Jodi Kearns, · 2014
Cited alongside, same era.
“LibrSspeech: An ASR corpus based on public domain audio books,”
“Purely sequence-trained neural networks for ASR based on lattice-free MMI,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahrmani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“BERT: Pre-training of deep bidirectional Transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Later among the works it cites.
“A time-restricted self-attention layer for ASR,”
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Listen, Attend and Spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
Jimmy Ba, Jamie Kiros, and Geoffrey Hinton, · 2016
Cited alongside, same era.
“On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition,”
Rohit Prabhavalkar, Ouais Alsharif, Antoine Bruguier, and Lan McGraw, · 2016
Cited alongside, same era.
“Model compression applied to small-footprint keyword spotting,”
George Tucker, Minhua Wu, Ming Sun, Sankaran Panchapagesan, Gengshen Fu, and Shiv Vitaladevuni, · 2016
Cited alongside, same era.
Later among the works it cites.
“Self-attention acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stuker, and Alex Waibel, · 2018
Later among the works it cites.
“Speech-Transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Later among the works it cites.
“Syllable-based sequence-to-sequence speech recognition with the Transformer in Mandarin Chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Later among the works it cites.
“Semi-orthogonal low-rank matrix factorization for deep neural networks,”
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohamadi, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“A pruned RNNLM lattice-rescoring algorithm for automatic speech recognition,”
Hainan Xu, Tongfei Chen, Dongji Gao, Yiming Wang, Ke Li, Nagendra Goel, Yishay Carmiel, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Capio 2017 conversational speech recognition system,”
Kyu J. Han, Akshay Chandrashekaran, Jungsuk Kim, and Ian R. Lane, · 2018
Later among the works it cites.
“A novel pyramidal-FSMN architecture with lattice-free MMI for speech recognition,”
Xuerui Yang, Jiwei Li, and Xi Zhou, · 2018
Later among the works it cites.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Closest in time.