Fetching the paper…
Reading the bibliography…
A key requirement in sequence to sequence processing is the modeling of long range dependencies.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Z. Yang, Y. Yang, W. W. Cohen, J. Carbonell, Q. V. Le, and R. Salakhutdinov 2019 · 1901
Earlier work this paper cites.
Guo, Q., X. Qiu, P. Liu, Y. Shao, X. Xue, and Z. Zhang 2019 · 1902
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., S. Gray, A. Radford, and I. Sutskever 2019 · 1904
Earlier work this paper cites.
The art of computer programming, vol. 3: Searching and sorting
Knuth, D. E. 1973 · 1973
Earlier work this paper cites.
A regular layout for parallel adders
Brent, R. P. and H. T. Kung 1982 · 1982
Earlier work this paper cites.
Sorting in c log n parallel steps
Ajtai, M., J. Komlós, and E. Szemerédi 1983 · 1983
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and J. Schmidhuber 1997 · 1997
Earlier work this paper cites.
Principles and practices of interconnection networks
Dally, W. J. and B. P. Towles 2004 · 2004
Earlier work this paper cites.
Sorting networks of logarithmic depth, further simplified
Seiferas, J. 2009 · 2009
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio 2014 · 2014
Earlier work this paper cites.
Graves, A., G. Wayne, and I. Danihelka 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and J. Ba 2014 · 2014
Earlier work this paper cites.
Learning to transduce with unbounded memory
Grefenstette, E., K. M. Hermann, M. Suleyman, and P. Blunsom 2015 · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A. and T. Mikolov 2015 · 2015
Earlier work this paper cites.
Kaiser, Ł. and I. Sutskever 2015 · 2015
Cited alongside, same era.
Kalchbrenner, N., I. Danihelka, and A. Graves 2015 · 2015
Cited alongside, same era.
Pointer networks
Vinyals, O., M. Fortunato, and N. Jaitly 2015 · 2015
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Yu, F. and V. Koltun 2015 · 2015
Cited alongside, same era.
Reinforcement learning neural Turing machines-revised
Zaremba, W. and I. Sutskever 2015 · 2015
Cited alongside, same era.
Simple and effective multi-paragraph reading comprehension
Clark, C. and M. Gardner 2017 · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Gehring, J., M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin 2017 · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin 2017 · 2017
Later among the works it cites.
Character-level language modeling with deeper self-attention
Al-Rfou, R., D. Choe, N. Constant, M. Guo, and L. Jones 2018 · 2018
Later among the works it cites.
Dehghani, M., S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrychowicz, M. and K. Kurach 2016 · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Graves, A., G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwińska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou, et al. 2016 · 2016
Cited alongside, same era.
Can active memory replace attention?
Kaiser, Ł. and S. Bengio 2016 · 2016
Cited alongside, same era.
Neural random access machines
Kurach, K., M. Andrychowicz, and I. Sutskever 2016 · 2016
Cited alongside, same era.
The lambada dataset: word prediction requiring a broad discourse context
Paperno, D., G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández 2016 · 2016
Cited alongside, same era.
Neural programmer-interpreters
Reed, S. E. and N. de Freitas 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
van den Oord, A., S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu 2016 · 2016
Cited alongside, same era.
Devlin, J., M.-W. Chang, K. Lee, and K. Toutanova 2018 · 2018
Later among the works it cites.
Improving the neural GPU architecture for algorithm learning
Freivalds, K. and R. Liepins 2018 · 2018
Later among the works it cites.
Recent advances in neural program synthesis
Kant, N. 2018 · 2018
Later among the works it cites.
Advances in pre-training distributed word representations
Mikolov, T., E. Grave, P. Bojanowski, C. Puhrsch, and A. Joulin 2018 · 2018
Later among the works it cites.
Divide and conquer networks
Nowak, A., D. Folqué, and J. Bruna 2018 · 2018
Later among the works it cites.
Tensor2tensor for neural machine translation
Vaswani, A., S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, Ł. Kaiser, N. Kalchbrenner, N. Parmar, et al. 2018 · 2018
Later among the works it cites.
Integer multiplication in time O(n log n)
Harvey, D. and J. Van Der Hoeven 2019 · 2019
Closest in time.
Music transformer
Huang, C.-Z. A., A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck 2019 · 2019
Closest in time.