Fetching the paper…
Reading the bibliography…
Tokenization is the first - and often underappreciated - layer of computation in language models.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt. 1958 · 1958
Earlier work this paper cites.
Polya’s theory of counting
Nicolaas Govert De Bruijn. 1964 · 1964
Earlier work this paper cites.
Counter machines and counter languages
Patrick C Fischer, Albert R Meyer, and Arnold L Rosenberg. 1968 · 1968
Earlier work this paper cites.
Children’s understanding of counting
Karen Wynn. 1990 · 1990
Earlier work this paper cites.
The computational complexity of counting
Mark Jerrum. 1995 · 1994
Earlier work this paper cites.
A recurrent neural network that learns to count
Paul Rodriguez, Janet Wiles, and Jeffrey L Elman. 1999 · 1999
Earlier work this paper cites.
Computability and logic
George S Boolos, John P Burgess, and Richard C Jeffrey. 2002 · 2002
Earlier work this paper cites.
Counter machines and verification problems
Oscar H Ibarra, Jianwen Su, Zhe Dang, Tevfik Bultan, and Richard A Kemmerer. 2002 · 2002
Earlier work this paper cites.
A survey of neural networks and formal languages
Joshua Ackerman and George Cybenko. 2020 · 2006
Earlier work this paper cites.
Deep autoregressive networks
Karol Gregor, Ivo Danihelka, Andriy Mnih, Charles Blundell, and Daan Wierstra. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich. 2015 · 2015
Cited alongside, same era.
Computability theory
S Barry Cooper. 2017 · 2017
Cited alongside, same era.
Attention is all you need
A Vaswani. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Cited alongside, same era.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Cited alongside, same era.
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li. 2023 · 2023
Later among the works it cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M. Shieber. 2019 · 2019
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, et al. 2022 · 2022
Cited alongside, same era.
A character-level length-control algorithm for non-autoregressive sentence summarization
Puyuan Liu, Xiang Zhang, and Lili Mou. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024a
Cited in the paper.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024b
Cited in the paper.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Language models need inductive biases to count inductively
Yingshan Chang and Yonatan Bisk. 2024 · 2024
Later among the works it cites.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2024 · 2024
Later among the works it cites.
Transformers, parallel computation, and logarithmic depth
Clayton Sanford, Daniel Hsu, and Matus Telgarsky. 2024 · 2024
Later among the works it cites.
Xiang Zhang, Muhammad Abdul-Mageed, and Laks V. S. Lakshmanan. 2024 · 2024
Later among the works it cites.
Xiang Zhang, Juntai Cao, Jiaqi Wei, Chenyu You, and Dujian Ding. 2025 · 2025
Closest in time.