Fetching the paper…
Reading the bibliography…
We introduce the Block-Recurrent Transformer, which applies a transformer layer in a recurrent fashion along a sequence, and has linear complexity with respect to sequence length.
R. J. Williams and J. Peng, “An efficient gradient-based algorithm for on-line training of recurrent network trajectories,” Neural Computation
1990
Earlier work this paper cites.
J. Schmidhuber, “Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets,” in International Conference on Artificial Neural Networks (ICANN)
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation
1997
Earlier work this paper cites.
R. K. Srivastava, K. Greff, and J. Schmidhuber, “Training very deep networks,” in NIPS
2015
Earlier work this paper cites.
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber, “LSTM: A search space odyssey,” IEEE transactions on neural networks and learning systems
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems
2017
Earlier work this paper cites.
U. Khandelwal, H. He, P. Qi, and D. Jurafsky, “Sharp nearby, fuzzy far away: How neural language models use context,” in Association for Computational Linguistics
2018
Earlier work this paper cites.
T. Lei, Y. Zhang, S. I. Wang, H. Dai, and Y. Artzi, “Simple recurrent units for highly parallelizable recurrence,” in EMNLP
2018
Earlier work this paper cites.
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. F. Foster, L. Jones, M. Schuster, N. Shazeer, N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, Z. Chen, Y. Wu, and M. Hughes, “The best of both worlds: Combining recent advances in neural machine translation,” in ACL
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in EMNLP
2018
Earlier work this paper cites.
N. Shazeer and M. Stern, “Adafactor: Adaptive learning rates with sublinear memory cost,” in ICML
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” in ACL
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontañón, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed, “Big bird: Transformers for longer sequences,” in NeurIPS
2020
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in International Conference on Learning Representations
2020
Earlier work this paper cites.
J. Ainslie, S. Ontañón, C. Alberti, V. Cvicek, Z. Fisher, P. Pham, A. Ravula, S. Sanghai, Q. Wang, and L. Yang, “ETC: encoding long and structured inputs in transformers,” in EMNLP
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. W. Rae, A. Potapenko, S. M. Jayakumar, C. Hillier, and T. P. Lillicrap, “Compressive transformers for long-range sequence modelling,” in ICLR
2020
Cited alongside, same era.
Z. Dai, G. Lai, Y. Yang, and Q. Le, “Funnel-transformer: Filtering out sequential redundancy for efficient language processing,” in NeurIPS
2020
Cited alongside, same era.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregressive transformers with linear attention,” in ICML
2020
Cited alongside, same era.
2020
Cited alongside, same era.
V. J. Hellendoorn, P. Maniatis, R. Singh, C. Sutton, and D. Bieber, “Global relational models of source code,” in ICLR
2020
K. Irie, I. Schlag, R. Csordás, and J. Schmidhuber, “Going beyond linear transformers with recurrent fast weight programmers,” in confNEU
2021
Later among the works it cites.
T. Lei, “When attention meets fast recurrence: Training language models with reduced compute,” in EMNLP
2021
Later among the works it cites.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in icml
2021
Later among the works it cites.
A. Al Adel and M. S. Burtsev, “Memory transformer with hierarchical attention for long document processing,” in 2021 International Conference Engineering and Telecommunication (En T)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research
2020
Cited alongside, same era.
A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen, “Query-key normalization for transformers,” in EMNLP
2020
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in NeurIPS
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler, “Long range arena : A benchmark for efficient transformers,” in International Conference on Learning Representations
2021
Cited alongside, same era.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics
2021
Cited alongside, same era.
2021
Later among the works it cites.
S. Sun, K. Krishna, A. Mattarella-Micke, and M. Iyyer, “Do long-range language models actually use long-range context?,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Later among the works it cites.
Y. Dong, J. Cordonnier, and A. Loukas, “Attention is not all you need: pure attention loses rank doubly exponentially with depth,” in ICML
2021
Later among the works it cites.
D. Hutchins, M. Rabe, Y. Wu, I. Schlag, and C. Staats, “Meliad.” Github source code repository. https://github.com/google-research/meliad , 2022
2022
Closest in time.
Y. Wu, M. Rabe, D. Hutchins, and C. Szegedy, “Memorizing transformers,” in ICLR
2022
Closest in time.
2022
Closest in time.
R. Csordás, K. Irie, and J. Schmidhuber, “The neural data router: Adaptive control flow in transformers improves systematic generalization,” in International Conference on Learning Representations
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
U. Shaham, E. Segal, M. Ivgi, A. Efrat, O. Yoran, A. Haviv, A. Gupta, W. Xiong, M. Geva, J. Berant, and O. Levy, “Scrolls: Standardized comparison over long language sequences,” 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
O. Press, N. Smith, and M. Lewis, “Train short, test long: Attention with linear biases enables input length extrapolation,” in International Conference on Learning Representations
2022
Closest in time.