Fetching the paper…
Reading the bibliography…
We investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition (ASR): hybrid Deep Bidirectional LSTM - Hidden Markov Models (DBLSTM-HMMs) and Recurrent Neural Network - Transducers (RNN-Ts).
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in 1992 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 1, 1992, pp. 517–520
1992
Earlier work this paper cites.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks,” in Proc. Interspeech , 2011, pp. 437–440
2011
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-rank matrix factorization for deep neural network training with high-dimensional output targets,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013, pp. 6655–6659
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 3586–3589
2015
Earlier work this paper cites.
J. H. Wong and M. J. Gales, “Sequence student-teacher training of deep neural networks,” in Proc. Interspeech , 2016, pp. 2761–2765
2016
Earlier work this paper cites.
Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.” in Proc. Interspeech , 2016, pp. 3439–3443
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim, B. Roomi, and P. Hall, “English conversational telephone speech recognition by humans and machines,” in Proc. Interspeech , 2017, pp. 132–136
2017
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6869–6898, 2017
2017
Earlier work this paper cites.
X. Cui, V. Goel, and G. Saon, “Embedding-based speaker adaptive training of deep neural networks,” in Proc. Interspeech , 2017, pp. 122–126
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
C. Xu, J. Yao, Z. Lin, W. Ou, Y. Cao, Z. Wang, and H. Zha, “Alternating multi-bit quantization for recurrent neural networks,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Li, R. Zhao, H. Hu, and Y. Gong, “Improving rnn transducer modeling for end-to-end speech recognition,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 114–121
2019
Later among the works it cites.
Y. You, J. Hseu, C. Ying, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large-batch training for lstm and beyond,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–16
2019
Later among the works it cites.
S. Venkataramani, J. Choi, V. Srinivasan, W. Wang, J. Zhang, M. Schaal, M. J. Serrano, K. Ishizaki, H. Inoue, E. Ogawa, M. Ohara, L. Chang, and K. Gopalakrishnan, “Deeptools: Compiler and execution runtime extensions for rapid ai accelerator,” IEEE Micro , vol. 39, no. 5, pp. 102–111, 2019
2019
Later among the works it cites.
S. Venkataramani, V. Srinivasan, J. Choi, P. Heidelberger, L. Chang, and K. Gopalakrishnan, “Memory and interconnect optimizations for peta-scale deep learning systems,” in 2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC) , 2019, pp. 225–234
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6381–6385
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Ardakani, Z. Ji, S. C. Smithson, B. H. Meyer, and W. J. Gross, “Learning recurrent binary/ternary weights,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
L. Hou, J. Zhu, J. Kwok, F. Gao, T. Qin, and T.-Y. Liu, “Normalization helps training of quantized lstm,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
W. Zhang, X. Cui, U. Finkler, G. Saon, A. Kayi, A. Buyuktosunoglu, B. Kingsbury, D. Kung, and M. Picheny, “A highly efficient distributed deep learning system for automatic speech recognition,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5706–5710
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-c. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. Interspeech , 2019, pp. 2613–2617
2019
Cited alongside, same era.
G. Saon, Z. Tüske, K. Audhkhasi, and B. Kingsbury, “Sequence noise injected training for end-to-end speech recognition,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6261–6265
2019
Cited alongside, same era.
2019
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single headed attention based sequence-to-sequence model for state-of-the-art results on switchboard-300,” in Proc. Interspeech , 2020, pp. 551–555
2020
Later among the works it cites.
2020
Later among the works it cites.
H. D. Nguyen, A. Alexandridis, and A. Mouchtaris, “Quantization aware training with absolute-cosine regularization for automatic speech recognition,” in Proc. Interspeech , 2020, pp. 3366–3370
2020
Later among the works it cites.
J. Xu, X. Chen, S. Hu, J. Yu, X. Liu, and H. Meng, “Low-bit quantization of recurrent neural network language models using alternating direction methods of multipliers,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7939–7943
2020
Later among the works it cites.
2021
Closest in time.
A. Agrawal, S. K. Lee, J. Silberman, M. Ziegler, M. Kang, S. Venkataramani, N. Cao, B. Fleischer, M. Guillorn, M. Cohen, S. Mueller, J. Oh, M. Lutz, J. Jung, S. Koswatta, C. Zhou, V. Zalani, J. Bonanno, R. Casatuta, C. Y. Chen, J. Choi, H. Haynie, A. Herbert, R. Jain, M. Kar, K. H. Kim, Y. Li, Z. Ren, S. Rider, M. Schaal, K. Schelm, M. Scheuermann, X. Sun, H. Tran, N. Wang, W. Wang, X. Zhang, V. Shah, B. Curran, V. Srinivasan, P. F. Lu, S. Shukla, L. Chang, and K. Gopalakrishnan, “9.1 a 7 nm 4-core ai chip with 25.6 tflops hybrid fp8 training, 102.4 tops int4 inference and workload-aware throttling,” in 2021 IEEE International Solid- State Circuits Conference (ISSCC) , vol. 64, 2021, pp. 144–146
2021
Closest in time.