Fetching the paper…
Reading the bibliography…
We explore options to use Transformer networks in neural transducer for end-to-end speech recognition.
“Class-based n-gram models of natural language,”
Peter F Brown, Peter V Desouza, Robert L Mercer, Vincent J Della Pietra, and Jenifer C Lai, · 1992
Earlier work this paper cites.
“Phoneme recognition using time-delay neural networks,”
Alexander Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang, · 1995
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Recurrent neural network based language model,”
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur, · 2010
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E Dahl, Dong Yu, Li Deng, and Alex Acero, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Achieving human parity in conversational speech recognition,”
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Convolutional sequence to sequence learning,”
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin, · 2017
Later among the works it cites.
“Automatic differentiation in pytorch,”
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, · 2017
Later among the works it cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Later among the works it cites.
Taku Kudo and John Richardson, · 2018
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Cited alongside, same era.
“Lower frame rate neural network acoustic models,”
Golan Pundak and Tara N Sainath, · 2016
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur Yi Li, Hairong Liu, Sanjeev Satheesh, Anuroop Sriram, and Zhenyao Zhu, · 2017
Cited alongside, same era.
Closest in time.
“Transformers with convolutional context for asr,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Closest in time.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Closest in time.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Closest in time.