Fetching the paper…
Reading the bibliography…
While Transformers have revolutionized deep learning, their quadratic attention complexity hinders their ability to process infinitely long inputs.
Brown, Tom u.a.(2020): Language models are few-shot learners
1901
Earlier work this paper cites.
Miller, George A(1956): The magical number seven, plus or minus two: Some limits on our capacity for processing information
1956
Earlier work this paper cites.
Wallace, Anthony FC (1960): Plans and the Structure of Behavior
1960
Earlier work this paper cites.
Fuster, Joaquin M(1973): Unit activity in prefrontal cortex during delayed-response performance: neuronal correlates of transient memory
1973
Earlier work this paper cites.
White, John G / Southgate, Eileen / Thomson, J Nichol(1986): S. Brenner (1986) The Structure of the Nervous System of the Nematode Caenorhabditis elegans
1986
Earlier work this paper cites.
LeCun, Yann / Bengio, Yoshua / others u.a.(1995): Convolutional networks for images, speech, and time series
1995
Earlier work this paper cites.
Hochreiter, Sepp / Schmidhuber, Jürgen(1997): Long Short-Term Memory
1997
Earlier work this paper cites.
Hochreiter, Sepp / Schmidhuber, Jürgen(1997): Long short-term memory
1997
Earlier work this paper cites.
Ashby, F Gregory / Ell, Shawn W / Valentin, Vivian V / Casale, Michael B(2005): FROST: A distributed neurocomputational model of working memory maintenance
2005
Earlier work this paper cites.
Baars, Bernard J(2005): Global workspace theory of consciousness: toward a cognitive neuroscience of human experience
2005
Earlier work this paper cites.
Cho, Kyunghyun / Van Merriënboer, Bart / Gulcehre, Caglar / Bahdanau, Dzmitry / Bougares, Fethi / Schwenk, Holger / Bengio, Yoshua(2014): Learning phrase representations using RNN encoder-decoder for statistical machine translation
2014
Earlier work this paper cites.
Tang, Xiaoyu / Wu, Jinglong / Shen, Yong(2016): The interactions of multisensory integration with endogenous and exogenous attention
2016
Earlier work this paper cites.
Chen, Tianqi / Xu, Bing / Zhang, Chiyuan / Guestrin, Carlos(2016): Training deep nets with sublinear memory cost
2016
Earlier work this paper cites.
Vaswani, Ashish / Shazeer, Noam / Parmar, Niki / Uszkoreit, Jakob / Jones, Llion / Gomez, Aidan N / Kaiser, Łukasz / Polosukhin, Illia(2017): Attention is all you need
2017
Earlier work this paper cites.
Christophel, Thomas B / Klink, P Christiaan / Spitzer, Bernhard / Roelfsema, Pieter R / Haynes, John Dylan(2017): The distributed nature of working memory
2017
Earlier work this paper cites.
Kirkpatrick, James u.a.(2017): Overcoming catastrophic forgetting in neural networks
2017
Earlier work this paper cites.
Shazeer, Noam / Stern, Mitchell(2018): Adafactor: Adaptive learning rates with sublinear memory cost
2018
Earlier work this paper cites.
Devlin, Jacob / Chang, Ming Wei / Lee, Kenton / Toutanova, Kristina(2018): Bert: Pre-training of deep bidirectional transformers for language understanding
2018
Earlier work this paper cites.
Hasani, Ramin / Lechner, Mathias / Amini, Alexander / Rus, Daniela / Grosu, Radu(2018): Can a Compact Neuronal Circuit Policy be Re-purposed to Learn Simple Robotic Control?
2018
Earlier work this paper cites.
Rangapuram, Syama Sundar / Seeger, Matthias W / Gasthaus, Jan / Stella, Lorenzo / Wang, Yuyang / Januschowski, Tim(2018): Deep state space models for time series forecasting
2018
Earlier work this paper cites.
Kočiskỳ, Tomáš / Schwarz, Jonathan / Blunsom, Phil / Dyer, Chris / Hermann, Karl Moritz / Melis, Gábor / Grefenstette, Edward(2018): The narrativeqa reading comprehension challenge
2018
Earlier work this paper cites.
Kudo, Taku / Richardson, John(2018): Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
2018
Earlier work this paper cites.
Rae, Jack W / Potapenko, Anna / Jayakumar, Siddhant M / Lillicrap, Timothy P(2019): Compressive transformers for long-range sequence modelling
2019
Earlier work this paper cites.
Shazeer, Noam(2019): Fast transformer decoding: One write-head is all you need
2019
Earlier work this paper cites.
Child, Rewon / Gray, Scott / Radford, Alec / Sutskever, Ilya(2019): Generating long sequences with sparse transformers
2019
Earlier work this paper cites.
Narayanan, Arun / Prabhavalkar, Rohit / Chiu, Chung Cheng / Rybach, David / Sainath, Tara N / Strohman, Trevor(2019): Recognizing long-form speech using streaming end-to-end models
2019
Earlier work this paper cites.
Guo, Qipeng / Qiu, Xipeng / Liu, Pengfei / Shao, Yunfan / Xue, Xiangyang / Zhang, Zheng(2019): Star-transformer
2019
Earlier work this paper cites.
Dai, Zihang / Yang, Zhilin / Yang, Yiming / Carbonell, Jaime / Le, Quoc V / Salakhutdinov, Ruslan(2019): Transformer-xl: Attentive language models beyond a fixed-length context
2019
Earlier work this paper cites.
Fan, Angela / Lavril, Thibaut / Grave, Edouard / Joulin, Armand / Sukhbaatar, Sainbayar(2020): Addressing some limitations of transformers with feedback memory
2020
Cited alongside, same era.
Zaheer, Manzil u.a.(2020): Big bird: Transformers for longer sequences
2020
Cited alongside, same era.
Gulati, Anmol u.a.(2020): Conformer: Convolution-augmented transformer for speech recognition
2020
Cited alongside, same era.
Ding, Siyu / Shang, Junyuan / Wang, Shuohuan / Sun, Yu / Tian, Hao / Wu, Hua / Wang, Haifeng(2020): ERNIE-Doc: A retrospective long-document modeling transformer
2020
Cited alongside, same era.
Gupta, Ankit / Berant, Jonathan(2020): Gmat: Global memory augmentation for transformers
2020
Cited alongside, same era.
Ramesh, Aditya / Dhariwal, Prafulla / Nichol, Alex / Chu, Casey / Chen, Mark(2022): Hierarchical text-conditional image generation with clip latents
2022
Later among the works it cites.
Wu, Yuhuai / Rabe, Markus N / Hutchins, DeLesley / Szegedy, Christian(2022): Memorizing transformers
2022
Later among the works it cites.
Bulatov, Aydar / Kuratov, Yury / Burtsev, Mikhail(2022): Recurrent memory transformer
2022
Later among the works it cites.
Chung, Hyung Won u.a.(2022): Scaling instruction-finetuned language models
2022
Later among the works it cites.
Tay, Yi u.a.(2022): Scaling laws vs model architectures: How does inductive bias influence scaling?
2022
Later among the works it cites.
Shaham, Uri u.a.(2022): Scrolls: Standardized comparison over long language sequences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dosovitskiy, Alexey u.a.(2020): An image is worth 16x16 words: Transformers for image recognition at scale
2020
Cited alongside, same era.
Wang, Sinong / Li, Belinda Z / Khabsa, Madian / Fang, Han / Ma, Hao(2020): Linformer: Self-attention with linear complexity
2020
Cited alongside, same era.
Beltagy, Iz / Peters, Matthew E / Cohan, Arman(2020): Longformer: The long-document transformer
2020
Cited alongside, same era.
Xiong, Ruibin u.a.(2020): On layer normalization in the transformer architecture
2020
Cited alongside, same era.
Kitaev, Nikita / Kaiser, Łukasz / Levskaya, Anselm(2020): Reformer: The efficient transformer
2020
Cited alongside, same era.
Choromanski, Krzysztof u.a.(2020): Rethinking attention with performers
2020
Cited alongside, same era.
Kaplan, Jared u.a.(2020): Scaling laws for neural language models
2020
Cited alongside, same era.
2022
Later among the works it cites.
Ju, Da / Roller, Stephen / Sukhbaatar, Sainbayar / Weston, Jason E(2022): Staircase attention for recurrent processing of sequences
2022
Later among the works it cites.
Villalobos, Pablo / Sevilla, Jaime / Heim, Lennart / Besiroglu, Tamay / Hobbhahn, Marius / Ho, Anson(2022): Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
2022
Later among the works it cites.
Chevalier, Alexis / Wettig, Alexander / Ajith, Anirudh / Chen, Danqi(2023): Adapting Language Models to Compress Contexts
2023
Later among the works it cites.
Borsos, Zalán u.a.(2023): Audiolm: a language modeling approach to audio generation
2023
Later among the works it cites.
Xiong, Wenhan u.a.(2023): Effective long-context scaling of foundation models
2023
Later among the works it cites.
Xiao, Guangxuan / Tian, Yuandong / Chen, Beidi / Han, Song / Lewis, Mike(2023): Efficient streaming language models with attention sinks
2023
Later among the works it cites.
Chen, Shouyuan / Wong, Sherman / Chen, Liangjian / Tian, Yuandong(2023): Extending context window of large language models via positional interpolation
2023
Later among the works it cites.
Tworkowski, Szymon / Staniszewski, Konrad / Pacek, Mikołaj / Wu, Yuhuai / Michalewski, Henryk / Miłoś, Piotr(2023): Focused transformer: Contrastive training for context scaling
2023
Later among the works it cites.
Team, Gemini u.a.(2023): Gemini: a family of highly capable multimodal models
2023
Later among the works it cites.
Achiam, Josh u.a.(2023): Gpt-4 technical report
2023
Later among the works it cites.
Mohtashami, Amirkeivan / Jaggi, Martin(2023): Landmark Attention: Random-Access Infinite Context Length for Transformers
2023
Later among the works it cites.
Mu, Jesse / Li, Xiang Lisa / Goodman, Noah(2023): Learning to compress prompts with gist tokens
2023
Later among the works it cites.
Gu, Albert / Dao, Tri(2023): Mamba: Linear-time sequence modeling with selective state spaces
2023
Later among the works it cites.
Jiang, Albert Q u.a.(2023): Mistral 7B
2023
Later among the works it cites.
Chowdhery, Aakanksha u.a.(2023): Palm: Scaling language modeling with pathways
2023
Later among the works it cites.
Peng, Bo u.a.(2023): RWKV: Reinventing RNNs for the Transformer Era
2023
Later among the works it cites.
Jouppi, Norm u.a.(2023): Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings
2023
Later among the works it cites.
Darcet, Timothée / Oquab, Maxime / Mairal, Julien / Bojanowski, Piotr(2023): Vision transformers need registers
2023
Later among the works it cites.
LMSYS (2023): LMSYS Chatbot Arena Leaderboard https://chat.lmsys.org/?arena
2023
Later among the works it cites.
Munkhdalai, Tsendsuren / Faruqui, Manaal / Gopal, Siddharth(2024): Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
2024
Closest in time.
Su, Jianlin / Ahmed, Murtadha / Lu, Yu / Pan, Shengfeng / Bo, Wen / Liu, Yunfeng(2024): Roformer: Enhanced transformer with rotary position embedding
2024
Closest in time.
Oren, Matanel / Hassid, Michael / Adi, Yossi / Schwartz, Roy(2024): Transformers are Multi-State RNNs
2024
Closest in time.