Fetching the paper…
Reading the bibliography…
A major challenge for transformers is generalizing to sequences longer than those observed during training.
On finite monoids having only trivial subgroups
Marcel Paul Schützenberger · 1965
Earlier work this paper cites.
Counter-Free Automata (MIT research monograph no. 65)
Robert McNaughton and Seymour A Papert · 1971
Earlier work this paper cites.
Dynamic construction of finite-state automata from examples using hill-climbing
Masaru Tomita · 1982
Earlier work this paper cites.
Regular languages in NC1
David A Mix Barrington, Kevin Compton, Howard Straubing, and Denis Thérien · 1992
Earlier work this paper cites.
Diamonds are forever: The variety DA
Pascal Tesson and Denis Thérien · 2002
Earlier work this paper cites.
Some results on majority quantifiers over words
K-J Lange · 2004
Earlier work this paper cites.
Linear circuits, two-variable logic and weakly blocked monoids
Christoph Behle, Andreas Krebs, and Mark Mercer · 2007
Earlier work this paper cites.
Typed semigroups, majority logic, and threshold circuits
Andreas Krebs · 2008
Earlier work this paper cites.
An Introduction to Kolmogorov Complexity and its Applications
Ming Li and Paul Vitányi · 2008
Earlier work this paper cites.
Regular languages definable by majority quantifiers with two variables
Christoph Behle, Andreas Krebs, and Stephanie Reifferscheid · 2009
Earlier work this paper cites.
Grammatical inference: learning automata and grammars
Colin De la Higuera · 2010
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
GLU variants improve transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Earlier work this paper cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Cited alongside, same era.
The regular languages of wire linear AC 0 {}^{\mbox{0}}
Michaël Cadilhac and Charles Paperman · 2022
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak · 2022
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Cited alongside, same era.
On provable length and compositional generalization
Kartik Ahuja and Amin Mansouri · 2024
Closest in time.
Logical languages accepted by transformer encoders with hard attention
Pablo Barcelo, Alexander Kozachinskiy, Anthony Widjaja Lin, and Vladimir Podolskii · 2024
Closest in time.
Separations in the representational capabilities of transformers and recurrent architectures
Satwik Bhattamishra, Michael Hahn, Phil Blunsom, and Varun Kanade · 2024
Closest in time.
Language models need inductive biases to count inductively
Yingshan Chang and Yonatan Bisk · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Cited alongside, same era.
Improving length-generalization in transformers via task hinting
Pranjal Awasthi and Anupam Gupta · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega · 2023
Cited alongside, same era.
Length generalization in arithmetic transformers
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2023
Cited alongside, same era.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Cited alongside, same era.
Closest in time.
Why are sensitive functions hard for transformers?
Michael Hahn and Mark Rofin · 2024
Closest in time.
Universal length generalization with turing programs
Kaiying Hou, David Brandfonbrener, Sham M. Kakade, Samy Jelassi, and Eran Malach · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M. Kakade, and Eran Malach · 2024
Closest in time.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2024
Closest in time.
Exposing attention glitches with flip-flop language modeling
Bingbin Liu, Jordan Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2024
Closest in time.
On limitations of the transformer architecture
Binghui Peng, Srini Narayanan, and Christos Papadimitriou · 2024
Closest in time.
One-layer transformers fail to solve the induction heads task
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
The expressive capacity of state space models: A formal language perspective
Yash Sarrof, Yana Veitsman, and Michael Hahn · 2024
Closest in time.
What Formal Languages Can Transformers Express? A Survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Length generalization of causal transformers without position encoding
Jie Wang, Tao Ji, Yuanbin Wu, Hang Yan, Tao Gui, Qi Zhang, Xuanjing Huang, and Xiaoling Wang · 2024
Closest in time.
Counting like transformers: Compiling temporal counting logic into softmax transformers
Andy Yang and David Chiang · 2024
Closest in time.
What algorithms can transformers learn? A study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Joshua M. Susskind, Samy Bengio, and Preetum Nakkiran · 2024
Closest in time.