Fetching the paper…
Reading the bibliography…
Transformer-based acoustic modeling has achieved great suc-cess for both hybrid and sequence-to-sequence speech recogni-tion.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu
2012
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in
2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka, “Neural Turing machines,”
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
A.-r. Mohamed, F. Seide, D. Yu, J. Droppo
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Earlier work this paper cites.
K. Chen and Q. Huo, “Training deep bidirectional LSTM acoustic model for LVCSR by a Context-Sensitive-Chunk BPTT approach,”
2016
Earlier work this paper cites.
Y. Zhang, G. Chen, D. Yu, K. Yao
2016
Earlier work this paper cites.
J. Lei Ba, J. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar
2017
Cited alongside, same era.
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, “Low latency acoustic modeling using temporal convolution and lstms,”
2017
Cited alongside, same era.
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina
2018
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. S
2018
2019
Later among the works it cites.
A. Mohamed, D. Okhonko, and L. Zettlemoyer, “Transformers with convolutional context for asr,”
2019
Later among the works it cites.
2019
Later among the works it cites.
O. Myle, E. Sergey, B. Alexei, F. Angela
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. Povey, H. Hadian, P. Ghahremani
2018
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang
2019
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi
2019
Cited alongside, same era.
Y. Wang, A. Mohamed, D. Le, C. Liu
2019
Cited alongside, same era.
L. Dong, F. Wang, and B. Xu, “Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,” in
2019
Cited alongside, same era.
E. Tsunoo, Y. Kashiwagi, T. Kumakura, and S. Watanabe, “Transformer ASR with Contextual Block Processing,” in
2019
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in
2020
Closest in time.
2020
Closest in time.
A. Tjandra, C. Liu, F. Zhang
2020
Closest in time.
Y. Shi, Y. Wang, C. Wu, C. Fuegen, F. Zhang, D. Le, C.-F. Yeh, and M. Seltzer, “Weak-Attention Suppression For Transformer Based Speech Recognition,” in
2020
Closest in time.