Fetching the paper…
Reading the bibliography…
The Transformer architecture is crucial for numerous AI models, but it still faces challenges in long-range language modeling.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney. 2011 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Recurrentgpt: Interactive generation of (arbitrarily) long text
Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, and Mrinmaya Sachan. 2023 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich Werlen, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019 · 2019
Cited alongside, same era.
When and why is document-level context useful in neural machine translation?
Yunsu Kim, Duc Thanh Tran, and Hermann Ney. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
Fast transformers with clustered attention
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020 · 2020
Later among the works it cites.
Krzysztof Choromanski, Haoxian Chen, Han Lin, Yuanzhe Ma, Arijit Sehanobish, Deepali Jain, Michael S Ryoo, Jake Varley, Andy Zeng, Valerii Likhosherstov, et al. 2021 · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Later among the works it cites.
Not all memories are created equal: Learning to forget by expiring
Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, and Angela Fan. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2019 · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Édouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Cited alongside, same era.
Smyrf-efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis. 2020 · 2020
Cited alongside, same era.
Abc: Attention with bounded-memory control
Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A Smith. 2022a
Cited in the paper.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2022b
Cited in the paper.
Linear complexity randomized self-attention mechanism
Lin Zheng, Chong Wang, and Lingpeng Kong. 2022a
Cited in the paper.
Sparsifying transformer models with trainable representation pooling
Michał Pietruszka, Łukasz Borchmann, and Łukasz Garncarek. 2022 · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew R Gormley. 2023 · 2023
Closest in time.