Fetching the paper…
Reading the bibliography…
We explore deep autoregressive Transformer models in language modeling for speech recognition.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmann machines,” in
2010
Earlier work this paper cites.
M. Sundermeyer, R. Schlüter, and H. Ney, “LSTM neural networks for language modeling.” in
2012
Earlier work this paper cites.
M. Sundermeyer, Z. Tüske, R. Schlüter, and H. Ney, “Lattice decoding and rescoring with long-span neural network language models,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
J. Cheng, L. Dong, and M. Lapata, “Long short-term memory-networks for machine reading,” in
2016
Earlier work this paper cites.
A. P. Parikh, O. Täckström, D. Das, and J. Uszkoreit, “A decomposable attention model for natural language inference,” in
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: a neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
M. Abadi
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
Z. Lin, M. Feng, C. N. d. Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, “A structured self-attentive sentence embedding,”
2017
Earlier work this paper cites.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” in
2017
Cited alongside, same era.
Ç. Gülçehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio, “On using monolingual corpora in neural machine translation,”
2017
Cited alongside, same era.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in
2017
Cited alongside, same era.
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber, “LSTM: A search space odyssey,”
2017
Cited alongside, same era.
A. Zeyer, P. Doetsch, P. Voigtlaender, R. Schlüter, and H. Ney, “A comprehensive study of deep bidirectional lstm rnns for acoustic modeling in speech recognition,” in
2017
Cited alongside, same era.
K. Irie, Z. Lei, R. Schlüter, and H. Ney, “Prediction of LSTM-RNN full context states as a subtask for N-gram feedforward language models,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, Ł. Kaiser, and N. Shazeer, “Generating wikipedia by summarizing long sequences,” in
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” [Online]. : https://blog.openai.com/language-unsupervised/, 2018
2018
Cited alongside, same era.
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, “Improved training of end-to-end attention models for speech recognition,” in
2018
Cited alongside, same era.
S. Toshniwal, A. Kannan, C.-C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A comparison of techniques for language model integration in encoder-decoder speech recognition,” in
2018
Cited alongside, same era.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in
2018
Cited alongside, same era.
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, “Self-attentional acoustic models,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Closest in time.
R. Al-Rfou, D. Choe, N. Constant, M. Guo, and L. Jones, “Character-level language modeling with deeper self-attention,” in
2019
Closest in time.
A. Baevski and M. Auli, “Adaptive input representations for neural language modeling,” in
2019
Closest in time.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” [Online]. : https://blog.openai.com/better-language-models/, 2019
2019
Closest in time.
J. Salazar, K. Kirchhoff, and Z. Huang, “Self-attention networks for connectionist temporal classification in speech recognition,” in
2019
Closest in time.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser, “Universal Transformers,” in
2019
Closest in time.
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “RWTH ASR systems for LibriSpeech: Hybrid vs Attention,” in
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.