Fetching the paper…
Reading the bibliography…
In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system.
Fundamentals of Speech Recognition
L. R. Rabiner and B.-H. Juang, · 1993
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Practical variational inference for neural networks,”
Alex Graves, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling,”
Haşim Sak, Andrew Senior, and Francoise Beaufays, · 2014
Earlier work this paper cites.
Navdeep Jaitly, David Sussillo, Quoc V Le, Oriol Vinyals, Ilya Sutskever, and Samy Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Streaming small-footprint keyword spotting using sequence-to-sequence models,”
Yanzhang (Ryan) He, Rohit Prabhavalkar, Kanishka Rao, Wei Li, Anton Bakhtin, and Ian McGraw, · 2017
Cited alongside, same era.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,”
Haşim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays, · 2017
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
in Proc. Interspeech 2018
“Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin chinese,” · 2018
Cited alongside, same era.
“A time-restricted self-attention layer for asr,”
“Very deep self-attention networks for end-to-end speech recognition,”
Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Müller, and Alex Waibel, · 2019
Later among the works it cites.
“Learning deep transformer models for machine translation,”
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao, · 2019
Later among the works it cites.
“Transformers with convolutional context for ASR,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Later among the works it cites.
“Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Later among the works it cites.
“Towards online end-to-end transformer automatic speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Self-attentional acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Cited alongside, same era.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Cited alongside, same era.
“Language Modeling with Deep Transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Cited alongside, same era.
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura, and Shinji Watanabe, · 2019
Later among the works it cites.
“Monotonic Recurrent Neural Network Transducer and Decoding Strategies,”
Anshuman Tripathi, Han Lu, Hasim Sak, and Hagen Soltau, · 2019
Later among the works it cites.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Duc Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, Christian Fuegen, Geoffrey Zweig, and Michael L. Seltzer, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.