Fetching the paper…
Reading the bibliography…
Empirical studies have identified a range of learnability biases and limitations of transformers, such as a persistent difficulty in learning to compute simple formal languages such as PARITY, and a bias towards low-degree functions.
The influence of variables on boolean functions
J. Kahn, G. Kalai, and N. Linial. 1988 · 1988
Earlier work this paper cites.
A brief introduction to fourier analysis on the boolean cube
Ronald De Wolf. 2008 · 2008
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2020 · 2010
Earlier work this paper cites.
Variations on the sensitivity conjecture
Pooya Hatami, Raghav Kulkarni, and Denis Pankratov. 2010 · 2010
Earlier work this paper cites.
Concise formulas for the area and volume of a hyperspherical cap
Shengqiao Li. 2010 · 2010
Earlier work this paper cites.
Boolean Function Complexity: Advances and Frontiers
Stasys Jukna. 2012 · 2012
Earlier work this paper cites.
Analysis of Boolean Functions
Ryan O’Donnell. 2014 · 2014
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville. 2019 · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2019 · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur*, Hossein Mobahi, Dilip Krishnan, and Samy Bengio. 2020 · 2020
Cited alongside, same era.
Sensitivity as a complexity measure for sequence classification tasks
Michael Hahn, Dan Jurafsky, and Richard Futrell. 2021 · 2021
Cited alongside, same era.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2021 · 2021
Cited alongside, same era.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Masked hard-attention transformers and boolean rasp recognize exactly the star-free languages
Dana Angluin, David Chiang, and Andy Yang. 2023 · 2023
Later among the works it cites.
Simplicity bias in transformers and their ability to learn sparse boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom. 2023 · 2023
Later among the works it cites.
Tighter bounds on the expressivity of transformer encoders
David Chiang, Peter Cholak, and Anand Pillay. 2023 · 2023
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee. 2023 · 2023
Later among the works it cites.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang. 2022 · 2022
Cited alongside, same era.
Spectral bias in practice: The role of function frequency in generalization
Sara Fridovich-Keil, Raphael Gontijo Lopes, and Rebecca Roelofs. 2022 · 2022
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A. Smith. 2022 · 2022
Cited alongside, same era.
On layer normalizations and residual connections in transformers
Sho Takase, Shun Kiyono, Sosuke Kobayashi, and Jun Suzuki. 2022 · 2022
Cited alongside, same era.
How does sharpness-aware minimization minimize sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li. 2022 · 2022
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2023 · 2023
Later among the works it cites.
On the maximum hessian eigenvalue and generalization
Simran Kaur, Jeremy Cohen, and Zachary Chase Lipton. 2023 · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak. 2023 · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023 · 2023
Later among the works it cites.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal. 2023a · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness. 2023 · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky. 2023 · 2023
Later among the works it cites.
Average-hard attention transformers are constant-depth uniform threshold circuits
Lena Strobl. 2023 · 2023
Later among the works it cites.
Transformers as recognizers of formal languages: A survey on expressivity
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin. 2023 · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran. 2023 · 2023
Later among the works it cites.