Fetching the paper…
Reading the bibliography…
Counting is a fundamental example of generalization, whether viewed through the mathematical lens of Peano's axioms defining the natural numbers or the cognitive science literature for children learning to count.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Children’s acquisition of the number words and the counting system
Karen Wynn · 1992
Earlier work this paper cites.
Rasp: A general logic synthesis system for sram-based fpgas
Jason Cong, John Peck, and Yuzheng Ding · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
How to learn the natural numbers: Inductive inference and the acquisition of number concepts
Eric Margolis and Stephen Laurence · 2008
Earlier work this paper cites.
How counting represents number: What children must learn and when they learn it
Barbara W Sarnecka and Susan Carey · 2008
Earlier work this paper cites.
Does learning to count involve a semantic induction?
Kathryn Davidson, Kortney Eng, and David Barner · 2012
Earlier work this paper cites.
Why neural translations are the right length
Xing Shi, Kevin Knight, and Deniz Yuret · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Meaning before order: Cardinal principle knowledge predicts improvement in understanding the successor principle and exact ordering
Elizabet Spaepen, Elizabeth A Gunderson, Dominic Gibson, Susan Goldin-Meadow, and Susan C Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Sequential neural networks as automata
William Merrill · 2019
Earlier work this paper cites.
LSTM networks can perform dynamic counting
Mirac Suzgun, Yonatan Belinkov, Stuart Shieber, and Sebastian Gehrmann · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov · 2019
Earlier work this paper cites.
On the ability of self-attention networks to recognize counter languages
S. Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
How can self-attention networks recognize dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Shape: Shifted absolute position embedding for transformers
Shun Kiyono, Sosuke Kobayashi, Jun Suzuki, and Kentaro Inui · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Cited alongside, same era.
Acquiring the cardinal knowledge of number words: A conceptual replication
Laurence Rousselle and Line Vossius · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Cited alongside, same era.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos H. Papadimitriou, and Karthik Narasimhan · 2021
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, J’anos Kram’ar, Matthew Rahtz, Tom McGrath, and Vladimir Mikulik · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Cited alongside, same era.
Simplicity bias in transformers and their ability to learn sparse boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak · 2022
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Gr’egoire Del’etang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A. Ortega · 2022
Cited alongside, same era.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Cited alongside, same era.
Anian Ruoss, Gr’egoire Del’etang, Tim Genewein, Jordi Grau-Moya, R. Csordás, Mehdi Abbana Bennani, Shane Legg, and Joel Veness · 2023
Later among the works it cites.
Testing the general deductive reasoning capacity of large language models using ood examples
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Seyed Mehran Kazemi, Najoung Kim, and He He · 2023
Later among the works it cites.
Efficient large language models: A survey
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, et al · 2023
Later among the works it cites.
Can transformers learn to solve problems recursively?
Shizhuo Zhang, Curt Tigges, Stella Biderman, Maxim Raginsky, and Talia Ringer · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2024
Closest in time.
Transformers can do arithmetic with the right embeddings
Sean McLeish, Arpit Bansal, Alex Stein, Neel Jain, John Kirchenbauer, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, et al · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Closest in time.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Michael Wornow, Callum Birch-Sykes, Stefano Massaroli, Aman Patel, Clayton Rabideau, Yoshua Bengio, et al · 2024
Closest in time.
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou, Przemys l aw Kazienko, G Kranthikiran, Jan Koco’n, Bartlomiej Koptyra, Satyapriya Krishna, Ronald McClelland, Niklas Muennighoff, Fares Obeid, Atsushi Saito, Guangyu Song, Haoqin Tu, Stanislaw Wo’zniak, Ruichong Zhang, Bingchen Zhao, Qihang Zhao, Peng Zhou, Jian Zhu, and Ruijie Zhu · 2024
Closest in time.
Lena Strobl, Dana Angluin, David Chiang, Jonathan Rawski, and Ashish Sabharwal · 2024
Closest in time.
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Tommaso Guidi, Marco Gori, and Stefano Melacci · 2024
Closest in time.
Counting like transformers: Compiling temporal counting logic into softmax transformers
Andy Yang and David Chiang · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Yongchao Zhou, Uri Alon, Xinyun Chen, Xuezhi Wang, Rishabh Agarwal, and Denny Zhou · 2024
Closest in time.