Fetching the paper…
Reading the bibliography…
Attention layers, as commonly used in transformers, form the backbone of modern deep learning, yet there is no mathematical description of their benefits and deficiencies as compared with other architectures.
Neighborly and cyclic polytopes
David Gale · 1963
Earlier work this paper cites.
The uniform convergence of frequencies of the appearance of events to their probabilities
Vladimir Naumovich Vapnik and Aleksei Yakovlevich Chervonenkis · 1968
Earlier work this paper cites.
On the density of families of sets
Norbert Sauer · 1972
Earlier work this paper cites.
A combinatorial problem; stability and order for models and theories in infinitary languages
Saharon Shelah · 1972
Earlier work this paper cites.
Some complexity questions related to distributive computing (preliminary report)
Andrew Chi-Chih Yao · 1979
Earlier work this paper cites.
Local and global properties in networks of processors
Dana Angluin · 1980
Earlier work this paper cites.
Parity, circuits, and the polynomial-time hierarchy
Merrick Furst, James B Saxe, and Michael Sipser · 1984
Earlier work this paper cites.
Monotone circuits for connectivity require super-logarithmic depth
Mauricio Karchmer and Avi Wigderson · 1988
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Distributed computing: a locality-sensitive approach
David Peleg · 2000
Earlier work this paper cites.
Limitations of learning via embeddings in euclidean half spaces
Shai Ben-David, Nadav Eiron, and Hans Ulrich Simon · 2002
Earlier work this paper cites.
Decoding by linear programming
Emmanuel J Candes and Terence Tao · 2005
Earlier work this paper cites.
Lectures on polytopes
Günter M Ziegler · 2006
Earlier work this paper cites.
Reconstruction and subgaussian operators in asymptotic geometric analysis
Shahar Mendelson, Alain Pajor, and Nicole Tomczak-Jaegermann · 2007
Earlier work this paper cites.
On the representational efficiency of restricted boltzmann machines
James Martens, Arkadev Chattopadhya, Toni Pitassi, and Richard Zemel · 2013
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Cited alongside, same era.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2018
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the equivalence between graph isomorphism testing and function approximation with GNNs
Zhengdao Chen, Soledad Villar, Lei Chen, and Joan Bruna · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Cited alongside, same era.
Universal invariant and equivariant graph neural networks
Nicolas Keriven and Gabriel Peyré · 2019
Cited alongside, same era.
What graph neural networks cannot learn: depth vs width
Andreas Loukas · 2019
Cited alongside, same era.
On the universality of invariant networks
Haggai Maron, Ethan Fetaya, Nimrod Segol, and Yaron Lipman · 2019
Cited alongside, same era.
Valerii Likhosherstov, Krzysztof Choromanski, and Adrian Weller · 2021
Later among the works it cites.
Size and depth separation in approximating benign functions with neural networks
Gal Vardi, Daniel Reichman, Toniann Pitassi, and Ohad Shamir · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos H. Papadimitriou, and Karthik Narasimhan · 2021
Later among the works it cites.
Exponentially improving the complexity of simulating the Weisfeiler-Lehman test with graph neural networks
Anders Aamand, Justin Chen, Piotr Indyk, Shyam Narayanan, Ronitt Rubinfeld, Nicholas Schiefer, Sandeep Silwal, and Tal Wagner · 2022
Later among the works it cites.
Simplicity bias in transformers and their ability to learn sparse boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Later among the works it cites.
Nuo Chen, Qiushi Sun, Renyu Zhu, Xiang Li, Xuesong Lu, and Ming Gao · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Later among the works it cites.
Exponential separations in symmetric neural networks
Aaron Zweig and Joan Bruna · 2022
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.