Fetching the paper…
Reading the bibliography…
This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition.
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
K. Vesely, A. Ghoshal, L. Burget, and D. Povey, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, and Others, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
R. Sennrich, B. Haddow, and A. Birch, · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Earlier work this paper cites.
“Self-attentional acoustic models,”
M. Sperber, J. Niehues, G. Neubig, and Others, · 2018
Earlier work this paper cites.
“Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin Chinese,”
S. Zhou, L. Dong, S. Xu, and B. Xu, · 2018
Earlier work this paper cites.
“A time-restricted self-attention layer for asr,”
D. Povey, H. Hadian, P. Ghahremani, and Others, · 2018
Earlier work this paper cites.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
T. Kudo and J. Richardson, · 2018
Cited alongside, same era.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Quoc V. Le, and R. Salakhutdinov, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,”
C. Raffel, N. Shazeer, A. Roberts, and Others, · 2019
Cited alongside, same era.
“A Comparative Study on Transformer vs RNN in Speech Applications,”
S. Karita, N. Chen, T. Hayashi, and Others, · 2019
Cited alongside, same era.
“Low Latency End-to-End Streaming Speech Recognition with a Scout Network,”
C. Wang, Y. Wu, S. Liu, J. Li, et al., · 2020
Closest in time.
“Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss,”
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, · 2020
Closest in time.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
A. Gulati, J. Qin, C.-C. Chiu, et al., · 2020
Closest in time.
“Fast, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces,”
F. Zhang, Y. Wang, X. Zhang, C. Liu, et al., · 2020
Closest in time.
“Streaming automatic speech recognition with the transformer model,”
N. Moritz, T. Hori, and J. L. Roux, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Transformer-Transducer: End-to-End Speech Recognition with Self-Attention,”
C.-F. Yeh, J. Mahadeokar, and Others, · 2019
Cited alongside, same era.
“Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition,”
J. Salazar, K. Kirchhoff, and Z. Huang, · 2019
Cited alongside, same era.
“Transformer-Based Acoustic Modeling for Hybrid Speech Recognition,”
Y. Wang, A. Mohamed, D. Le, and Others, · 2019
Cited alongside, same era.
“Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
L. Dong, F. Wang, and B. Xu, · 2019
Cited alongside, same era.
“From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition,”
D. Le, X. Zhang, W. Zheng, and Others, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D S Park, W Chan, Y Zhang, and Others, · 2019
Cited alongside, same era.
C. Wu, Y. Shi, Y. Wang, and C.-F. Yeh, · 2020
Closest in time.
“Exploring Transformers for Large-Scale Speech Recognition,”
L. Lu, C. Liu, J. Li, and Y. Gong, · 2020
Closest in time.
Y. Wang, Y. Shi, F. Zhang, C. Wu, and Others, · 2020
Closest in time.
“Weak-Attention Suppression For Transformer Based Speech Recognition,”
Y. Shi, Y. Wang, C. Wu, C. Fuegen, et al., · 2020
Closest in time.
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2020
Closest in time.
“Deja-vu: Double Feature Presentation and Iterated loss in Deep Transformer Networks,”
A. Tjandra, C. Liu, F. Zhang, and Others, · 2020
Closest in time.