Fetching the paper…
Reading the bibliography…
Recent research in the field of machine learning has increasingly focused on the memorization capacity of Transformers, but how efficient they are is not yet well understood.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M. Cover · 1965
Earlier work this paper cites.
Learning Machines
Nils J. Nilsson · 1965
Earlier work this paper cites.
Perceptrons: An Introduction to Computational Geometry
Marvin Minsky and Seymour Papert · 1969
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B Baum · 1988
Earlier work this paper cites.
Bounding the vapnik-chervonenkis dimension of concept classes parameterized by real numbers
Paul W. Goldberg and Mark R. Jerrum · 1995
Earlier work this paper cites.
Shattering All Sets of \CJK@punctchar
Eduardo D. Sontag · 1997
Earlier work this paper cites.
Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions
Guang-Bin Huang and H.A. Babri · 1998
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep Sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Optuna: A next-generation hyperparameter optimization framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama · 2019
Earlier work this paper cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Are Transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
MetaFormer is Actually What You Need for Vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean Functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2023
Later among the works it cites.
Vision Transformers Need Registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2023
Later among the works it cites.
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
Tokio Kajitsuka and Issei Sato · 2023
Later among the works it cites.
Provable Memorization Capacity of Transformers
Junghwan Kim, Michelle Kim, and Barzan Mozafari · 2023
Later among the works it cites.
Memorization Capacity of Multi-Head Attention in Transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O(n) Connections are Expressive Enough: Universal Approximability of Sparse Transformers
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Cited alongside, same era.
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Deep double descent: where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Provable Memorization via Deep Neural Networks using Sub-linear Parameters
Sejun Park, Jaeho Lee, Chulhee Yun, and Jinwoo Shin · 2021
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Inductive Biases and Variable Creation in Self-Attention Mechanisms
Benjamin L. Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Cited alongside, same era.
Later among the works it cites.
Scalable Diffusion Models with Transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
Representational Strengths and Limitations of Transformers, June 2023
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Later among the works it cites.
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
Shokichi Takakura and Taiji Suzuki · 2023
Later among the works it cites.
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
Xingwu Chen and Difan Zou · 2024
Closest in time.
Approximation Rate of the Transformer Architecture for Sequence Modeling, February 2024
Haotian Jiang and Qianxiao Li · 2024
Closest in time.
Liam Madden, Curtis Fox, and Christos Thrampoulidis · 2024
Closest in time.
Jonathan W. Siegel · 2024
Closest in time.
Sequence Length Independent Norm-Based Generalization Bounds for Transformers
Jacob Trauger and Ambuj Tewari · 2024
Closest in time.