Fetching the paper…
Reading the bibliography…
While significant improvements have been made in recent years in terms of end-to-end automatic speech recognition (ASR) performance, such improvements were obtained through the use of very large neural networks, unfit for embedded use on edge devices.
“Optimal brain damage,”
Yann LeCun, John S Denker, and Sara A Solla, · 1990
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Adadelta: an adaptive learning rate method,”
M. D. Zeiler, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Learning phrase representations using RNN encoder–decoder for statistical machine translation,”
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Sequence-level knowledge distillation,”
Yoon Kim and Alexander M Rush, · 2016
Cited alongside, same era.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al., · 2016
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Model compression via distillation and quantization,”
A. Polino, R. Pascanu, and D. Alistarh, · 2018
Later among the works it cites.
T. Kudo and J. Richardson, · 2018
Later among the works it cites.
“Transformers with convolutional context for asr,”
A. Mohamed, D. Okhonko, and L. Zettlemoyer, · 2019
Closest in time.
“A comparative study on transformer vs rnn in speech applications,”
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, et al., · 2019
Closest in time.
“Knowledge distillation using output errors for self-attention end-to-end models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Image transformer,”
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, · 2018
Cited alongside, same era.
“Quantization and training of neural networks for efficient integer-arithmetic-only inference,”
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices.,”
Y. He, T. N. Sainath, R. Prabhavalkar, I. Mcgraw, R. Alvarez, D. Zhao, D. Rybach, Y. Kannan, A. Wu, and R et al. Pang, · 2018
Cited alongside, same era.
“An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”
Shaojie Bai, J Zico Kolter, and Vladlen Koltun, · 2018
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever,
Cited in the paper.
H.-G. Kim, H. Na, H. Lee, J. Lee, T. G. Kang, M.-J. Lee, and Y. S. Choi, · 2019
Closest in time.
“Fully quantized transformer for improved translation,”
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh, · 2019
Closest in time.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Closest in time.
“Language modeling with deep transformers,”
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, · 2019
Closest in time.
“Efficient 8-bit quantization of transformer neural machine language translation model,”
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore, · 2019
Closest in time.
“Q8bert: Quantized 8bit bert,”
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat, · 2019
Closest in time.