Fetching the paper…
Reading the bibliography…
We present the Compressive Transformer, an attentive sequence model which compresses past memories for long-range sequence learning.
Dynamic evaluation of transformer language models
B. Krause, E. Kahembwe, I. Murray, and S. Renals · 1904
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Latent dirichlet allocation
D. M. Blei, A. Y. Ng, and M. I. Jordan · 2003
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
The human knowledge compression contest
M. Hutter · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
The goldilocks principle: Reading children’s books with explicit memory representations
F. Hill, A. Bordes, S. Chopra, and J. Weston · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Earlier work this paper cites.
Quasi-recurrent neural networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2016
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
J. Chung, S. Ahn, and Y. Bengio · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwińska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou, et al · 2016
Cited alongside, same era.
D. Ha, A. Dai, and Q. V. Le · 2016
Cited alongside, same era.
Neural machine translation in linear time
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. v. d. Oord, A. Graves, and K. Kavukcuoglu · 2016
Parallel wavenet: Fast high-fidelity speech synthesis
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg, et al · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
J. W. Rae, C. Dyer, P. Dayan, and T. P. Lillicrap · 2018
Later among the works it cites.
Relational recurrent neural networks
A. Santoro, R. Faulkner, D. Raposo, J. Rae, M. Chrzanowski, T. Weber, D. Wierstra, O. Vinyals, R. Pascanu, and T. Lillicrap · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
S. Smith, P. jan Kindermans, C. Ying, and Q. V. Le · 2018
Later among the works it cites.
End-to-end dense video captioning with masked transformer
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong · 2018
Later among the works it cites.
Character-level language modeling with deeper self-attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multiplicative lstm for sequence modelling
B. Krause, L. Lu, I. Murray, and S. Renals · 2016
Cited alongside, same era.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, R. Fernández, K. Erk, et al · 2016
Cited alongside, same era.
Scaling memory-augmented neural networks with sparse reads and writes
J. Rae, J. J. Hunt, I. Danihelka, T. Harley, A. W. Senior, G. Wayne, A. Graves, and T. Lillicrap · 2016
Cited alongside, same era.
The persistence and transience of memory
B. A. Richards and P. W. Frankland · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
R. Al-Rfou, D. Choe, N. Constant, M. Guo, and L. Jones · 2019
Closest in time.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Closest in time.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Closest in time.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, W. W. Cohen, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Closest in time.
The curious case of neural text degeneration
A. Holtzman, J. Buys, M. Forbes, and Y. Choi · 2019
Closest in time.
Large memory layers with product keys
G. Lample, A. Sablayrolles, M. Ranzato, L. Denoyer, and H. Jégou · 2019
Closest in time.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Closest in time.
Adaptive attention span in transformers
S. Sukhbaatar, E. Grave, P. Bojanowski, and A. Joulin · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli · 2019
Closest in time.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Closest in time.