Fetching the paper…
Reading the bibliography…
Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities.
M. Ostendorf, “Moving beyond the ‘beads-on-a-string’model of speech,” in
1999
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, and Others, “The Kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
N. Kalchbrenner and P. Blunsom, “Recurrent Continuous Translation Models,” in
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in
2014
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in
2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. D. Manning, “Effective Approaches to Attention-based Neural Machine Translation,” in
2015
Cited alongside, same era.
R. Jozefowicz, W. Zaremba, and I. Sutskever, “An Empirical Exploration of Recurrent Network Architectures,” in
2015
Cited alongside, same era.
J. Cheng, L. Dong, and M. Lapata, “Long short-term memory-networks for machine reading,” in
2016
Cited alongside, same era.
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit, “A Decomposable Attention Model for Natural Language Inference,” in
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
Z. Lin, M. Feng, C. N. dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, “A Structured Self-attentive Sentence Embedding,” in
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in
2017
Later among the works it cites.
Y. Zhang, W. Chan, and N. Jaitly, “Very Deep Convolutional Networks for End-to-End Speech Recognition,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Im and S. Cho, “Distance-based Self-Attention Network for Natural Language Inference,”
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Gal and Z. Ghahramani, “A Theoretically Grounded Application of Dropout in Recurrent Neural Networks,” in
2016
Cited alongside, same era.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” in
2016
Cited alongside, same era.
D. Povey, H. Hadian, P. Ghahremani, K. Li, and S. Khudanpur, “A Time-Restricted Self-Attention Layer for ASR,” in
2018
Closest in time.
G. Neubig, M. Sperber, X. Wang, M. Felix, A. Matthews, S. Padmanabhan, Y. Qi, D. S. Sachan, P. Arthur, P. Godard, J. Hewitt, R. Riad, and L. Wang, “XNMT: The eXtensible Neural Machine Translation Toolkit,” in
2018
Closest in time.
T. Q. Nguyen and D. Chiang, “Improving Lexical Choice in Neural Machine Translation,” in
2018
Closest in time.