Fetching the paper…
Reading the bibliography…
Reducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios.
2012
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large Scale Deep Neural Network Acoustic Modeling with Semi-supervised Training Data for YouTube Video Transcription,” in Proc. ASRU , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in ICONIP , 2015, pp. 577–585
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
G. Pundak and T. N. Sainath, “Lower frame rate neural network acoustic models,” in Proc. Interspeech , 2016
2016
Earlier work this paper cites.
R. Alvarez, R. Prabhavalkar, and A. Bakhtin, “On the efficient representation and execution of deep acoustic models,” Interspeech 2016 , pp. 2746–2750, 2016
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , 2017
2017
Earlier work this paper cites.
R. Takeda, K. Nakadai, and K. Komatani, “Node pruning based on entropy of weights and node activity for small-footprint acoustic model based on deep neural networks.” in INTERSPEECH , 2017, pp. 1636–1640
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Earlier work this paper cites.
C.-C. Chiu, T. N. Sainath, Y. Wu et al. , “State-of-the-art Speech Recognition With Sequence-to-Sequence Models,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
C. Li, L. Zhu, S. Xu, P. Gao, and B. Xu, “Compression of acoustic model via knowledge distillation and pruning,” in ICPR . IEEE, 2018, pp. 2785–2790
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D. Wang, X. Wang, and S. Lv, “An overview of end-to-end automatic speech recognition,” Symmetry , vol. 11, no. 8, p. 1018, 2019
2019
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6381–6385
2019
Cited alongside, same era.
D. Gao, X. He, Z. Zhou, Y. Tong, K. Xu, and L. Thiele, “Rethinking pruning for accelerating deep inference at the edge,” in SIGKDD , 2020, pp. 155–164
2020
Later among the works it cites.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S.-y. Chang, W. Li, R. Alvarez, Z. Chen et al. , “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6059–6063
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” Proc. Interspeech 2020 , pp. 5036–5040, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Li, R. Zhao, H. Hu, and Y. Gong, “Improving RNN transducer modeling for end-to-end speech recognition,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, D. Rybach, T. N. Sainath, and T. Strohman, “Recognizing long-form speech using streaming end-to-end models,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 920–927
2019
Cited alongside, same era.
J. Shen, P. Nguyen, Y. Wu, Z. Chen et al. , “Lingvo: a modular and scalable framework for sequence-to-sequence modeling,” 2019
2019
Cited alongside, same era.
J. Li, Y. Wu, Y. Gaur et al. , “On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A new training pipeline for an improved neural transducer,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
[Online]. Available: https://www.tensorflow.org/lite/performance/post_training_quantization
Cited in the paper.
2020
Later among the works it cites.
H. D. Nguyen, A. Alexandridis, and A. Mouchtaris, “Quantization aware training with absolute-cosine regularization for automatic speech recognition.” in Interspeech , 2020, pp. 3366–3370
2020
Later among the works it cites.
S. Ding, T. Chen, and Z. Wang, “Audio lottery: Speech recognition made ultra-lightweight, noise-robust, and transferable,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
A. Abdolrashidi, L. Wang, S. Agrawal, J. Malmaud, O. Rybakov, C. Leichner, and L. Lew, “Pareto-optimal quantized resnet is mostly 4-bit,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 3091–3099
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
T. N. Sainath, Y. He, A. Narayanan et al. , “An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling,” in Proc. of Interspeech , 2021
2021
Later among the works it cites.
R. Botros, T. Sainath, R. David, E. Guzman, W. Li, and Y. He, “Tied & reduced rnn-t decoder,” in Proc. Interspeech , 2021
2021
Later among the works it cites.