Fetching the paper…
Reading the bibliography…
Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory.
Augmenting self-attention with persistent memory
Sukhbaatar, S., Grave, E., Lample, G., Jegou, H., and Joulin, A · 1907
Earlier work this paper cites.
Finding structure in time
Elman, J · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Fan, A., Lavril, T., Grave, E., Joulin, A., and Sukhbaatar, S · 2002
Earlier work this paper cites.
The psychology and neuroscience of forgetting
Wixted, J. T · 2004
Earlier work this paper cites.
Fifty years of memory of college grades: Accuracy and distortions
Bahrick, H. P., Hall, L. K., and Da Costa, L. A · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Large text compression benchmark
Mahoney, M · 2011
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Replication and analysis of ebbinghaus’ forgetting curve
Murre, J. and Dros, J · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Efficient softmax approximation for gpus
Grave, E., Joulin, A., Cissé, M., and Jégou, H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Wu, F., Fan, A., Baevski, A., Dauphin, Y., and Auli, M · 2018
Cited alongside, same era.
Improving the transformer translation model with document-level context
Zhang, J., Luan, H., Sun, M., Zhai, F., Xu, J., Zhang, M., and Liu, Y · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Cited alongside, same era.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Cited alongside, same era.
Bp-transformer: Modelling long-range context via binary partitioning
Ye, Z., Guo, Q., Gan, Q., Qiu, X., and Zhang, Z · 2019
Later among the works it cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Davis, J., Sarlos, T., Belanger, D., Colwell, L., and Weller, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Adaptively sparse transformers
Correia, G. M., Niculae, V., and Martins, A. F · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Using local knowledge graph construction to scale seq2seq models to multi-document inputs
Fan, A., Gardent, C., Braud, C., and Bordes, A · 2019
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2019
Cited alongside, same era.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
Cited alongside, same era.
Izacard, G. and Grave, E · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gülçehre, Ç., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., Botvinick, M., Heess, N., and Hadsell, R · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Later among the works it cites.
Recipes for building an open-domain chatbot
Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Shuster, K., Smith, E. M., et al · 2020
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Later among the works it cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Closest in time.
Linear transformers are secretly fast weight memory systems
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Closest in time.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Closest in time.