Fetching the paper…
Reading the bibliography…
We conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory.
The theory of approximation , volume 11
Dunham Jackson · 1930
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
Brown corpus manual
W Nelson Francis and Henry Kucera · 1979
Earlier work this paper cites.
Neural net approximation
Andrew R Barron · 1992
Earlier work this paper cites.
Wavelets and Operators: Volume 1
Yves Meyer · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Sentiment analysis: Capturing favorability using natural language processing
Tetsuya Nasukawa and Jeonghee Yi · 2003
Earlier work this paper cites.
Deterministic dependency parsing of english text
Joakim Nivre and Mario Scholz · 2004
Earlier work this paper cites.
Compressed sensing
David L Donoho · 2006
Earlier work this paper cites.
An introduction to compressive sampling
Emmanuel J Candès and Michael B Wakin · 2008
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Earlier work this paper cites.
A priori estimates of the population risk for two-layer neural networks
Weinan E, Chao Ma, and Lei Wu · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Earlier work this paper cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Chao Ma, Stephan Wojtowytsch, Lei Wu, and Weinan E · 2020
Why self-attention is natural for sequence-to-sequence problems? a perspective from symmetries
Chao Ma and Lexing Ying · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Making transformers solve compositional tasks
Santiago Ontanón, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
The barron space and the flow-induced function spaces for neural network models
Weinan E, Chao Ma, and Lei Wu · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
A thousand brains: A new theory of intelligence
Jeff Hawkins · 2021
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al · 2022
Later among the works it cites.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Later among the works it cites.
Transformers learn through gradual rank increase
Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe, Samy Bengio, and Joshua Susskind · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Later among the works it cites.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Yuntian Gu, Bohang Zhang, Haotian Ye, Di He, and Liwei Wang · 2023
Later among the works it cites.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Later among the works it cites.
Approximation theory of transformer networks for sequence modeling
Haotian Jiang and Qianxiao Li · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2023
Later among the works it cites.
The lazy neuron phenomenon: On emergence of activation sparsity in transformers
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al · 2023
Later among the works it cites.
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Do pretrained transformers really learn in-context by gradient descent?
Lingfeng Shen, Aayush Mishra, and Daniel Khashabi · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Max Vladymyrov, Razvan Pascanu, et al · 2023
Later among the works it cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun · 2023
Later among the works it cites.
Understanding multi-phase optimization dynamics and rich nonlinear behaviors of relu networks
Mingze Wang and Chao Ma · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Anchor function: a type of benchmark functions for studying language models
Zhongwang Zhang, Zhiwei Wang, Junjie Yao, Zhangchen Zhou, Xiaolong Li, Zhi-Qin John Xu, et al · 2024
Closest in time.