Fetching the paper…
Reading the bibliography…
Unneeded elements in the attention's context degrade performance.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 1905
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Root mean square layer normalization, 2019
Biao Zhang and Rico Sennrich · 1910
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 1911
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling, 2019
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap · 1911
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 1911
Earlier work this paper cites.
Learning internal representations by error propagation, 1986
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2002
Earlier work this paper cites.
On layer normalization in the transformer architecture, 2020
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu · 2002
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2004
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2006
Earlier work this paper cites.
Big bird: Transformers for longer sequences, 2021
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed · 2007
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling, 2014
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations, 2016
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Costs of selective attention: When children notice what adults miss
Daniel J. Plebanek and Vladimir M. Sloutsky · 2017
Cited alongside, same era.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Cited alongside, same era.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering, 2018
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Primer: Searching for efficient transformers for language modeling, 2022
David R. So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V. Le · 2022
Later among the works it cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Later among the works it cites.
Optimizing retrieval-augmented reader models via token elimination, 2023
Moshe Berchansky, Peter Izsak, Avi Caciularu, Ido Dagan, and Moshe Wasserblat · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters, 2023
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Patrick Collier, Alexey Gritsenko, Vighnesh Birodkar, Cristina Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov, Filip Pavetić, Dustin Tran, Thomas Kipf, Mario Lučić, Xiaohua Zhai, Daniel Keysers, Jeremiah Harmsen, and Neil Houlsby · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Cited alongside, same era.
Multi-passage BERT: A globally normalized BERT model for open-domain question answering
Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, and Bing Xiang · 2019
Cited alongside, same era.
Sparse is enough in scaling transformers, 2021
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Cited alongside, same era.
Linear transformers are secretly fast weight programmers, 2021
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces, 2022
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
Later among the works it cites.
Longnet: Scaling transformers to 1,000,000,000 tokens, 2023
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, Nanning Zheng, and Furu Wei · 2023
Later among the works it cites.
Context compression for auto-regressive transformers with sentinel tokens, 2023
Siyu Ren, Qi Jia, and Kenny Q. Zhu · 2023
Later among the works it cites.
Focus on the core: Efficient attention via pruned token compression for document classification
Jungmin Yun, Mihyeon Kim, and Youngbin Kim · 2023
Later among the works it cites.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization, 2023
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
Dynamic context pruning for efficient and interpretable autoregressive transformers, 2024
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurelien Lucchi, and Thomas Hofmann · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
Model tells you what to discard: Adaptive kv cache compression for llms, 2024
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Albert Gu and Tri Dao · 2024
Closest in time.
Learning to compress prompts with gist tokens, 2024
Jesse Mu, Xiang Lisa Li, and Noah Goodman · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention, 2024
Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal · 2024
Closest in time.
Transformers are multi-state rnns, 2024
Matanel Oren, Michael Hassid, Nir Yarden, Yossi Adi, and Roy Schwartz · 2024
Closest in time.
Efficient attention: Attention with linear complexities, 2024
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li · 2024
Closest in time.