Fetching the paper…
Reading the bibliography…
Algorithmic reasoning requires capabilities which are most naturally understood through recurrent models of computation, like the Turing machine.
Produit complet des groupes de permutations et probleme d’extension de groupes II
Marc Krasner and Léo Kaloujnine · 1951
Earlier work this paper cites.
The algebraic theory of context-free languages
Noam Chomsky and Marcel P Schützenberger · 1959
Earlier work this paper cites.
Algebraic theory of machines, I: Prime decomposition theorem for finite semigroups and machines
Kenneth Krohn and John Rhodes · 1965
Earlier work this paper cites.
On finite monoids having only trivial subgroups
Marcel Paul Schützenberger · 1965
Earlier work this paper cites.
Cascade synthesis of finite-state machines
H Paul Zeiger · 1967
Earlier work this paper cites.
Automata, languages, and machines
Samuel Eilenberg · 1974
Earlier work this paper cites.
The number of semigroups of order n
Daniel J Kleitman, Bruce R Rothschild, and Joel H Spencer · 1976
Earlier work this paper cites.
Unbounded fan-in circuits and associative functions
Ashok K Chandra, Steven Fortune, and Richard Lipton · 1983
Earlier work this paper cites.
Parity, circuits, and the polynomial-time hierarchy
Merrick Furst, James B. Saxe, and Michael Sipser · 1984
Earlier work this paper cites.
Bounded-width polynomial-size branching programs recognize exactly those languages in 𝖭𝖢 1 \mathsf{NC}^{1}
David A. Mix Barrington · 1986
Earlier work this paper cites.
Data parallel algorithms
W Daniel Hillis and Guy L. Steele Jr · 1986
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Finite monoids and the fine structure of 𝖭𝖢 1 \mathsf{NC}^{1}
David A. Mix Barrington and Denis Thérien · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Finite permutation groups with large abelian quotients
László Kovács and Cheryl Praeger · 1989
Earlier work this paper cites.
Finite-automaton aperiodicity is PSPACE-complete
Sang Cho and Dung T Huynh · 1991
Earlier work this paper cites.
On threshold circuits and polynomial computation
John H. Reif and Stephen R. Tate · 1992
Earlier work this paper cites.
On the computational power of neural nets
Hava T Siegelmann and Eduardo D Sontag · 1992
Earlier work this paper cites.
On the cascaded decomposition of automata, its complexity and its application to logic (Draft)
Oded Maler and Amir Pnueli · 1994
Earlier work this paper cites.
Threshold circuits for iterated matrix product and powering
Carlo Mereghetti and Beatrice Palano · 2000
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
On the Krohn-Rhodes cascaded decomposition theorem
Oded Maler · 2010
Earlier work this paper cites.
Applications of automata theory and algebra: via the mathematical theory of complexity to biology, physics, psychology, philosophy, and games
John Rhodes, Chrystopher L Nehaniv, and Morris W Hirsch · 2010
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Computational holonomy decomposition of transformation semigroups
Attila Egri-Nagy and Chrystopher L Nehaniv · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
FractalNet: Ultra-deep neural networks without residuals
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
Reliably learning the ReLU in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2017
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher · 2017
Cited alongside, same era.
On the ability of neural nets to express distributions
Holden Lee, Rong Ge, Tengyu Ma, Andrej Risteski, and Sanjeev Arora · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Lower bounds over boolean inputs for deep neural networks with ReLU gates
Towards lower bounds on the depth of ReLU neural networks
Christoph Hertrich, Amitabh Basu, Marco Di Summa, and Martin Skutella · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Finetuning pretrained transformers into rnns
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, and Noah A Smith · 2021
Later among the works it cites.
On the power of saturated Transformers: A view from circuit complexity
William Merrill, Yoav Goldberg, Roy Schwartz, and Noah A. Smith · 2021
Later among the works it cites.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anirbit Mukherjee and Amitabh Basu · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Cited alongside, same era.
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Later among the works it cites.
Attention is turing complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Later among the works it cites.
Can contrastive learning avoid shortcut solutions?
Joshua Robinson, Li Sun, Ke Yu, Kayhan Batmanghelich, Stefanie Jegelka, and Suvrit Sra · 2021
Later among the works it cites.
Programming puzzles
Tal Schuster, Ashwin Kalyan, Alex Polozov, and Adam Kalai · 2021
Later among the works it cites.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein · 2021
Later among the works it cites.
Scaling local self-attention for parameter efficient visual backbones
Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake A. Hechtman, and Jonathon Shlens · 2021
Later among the works it cites.
Thinking like Transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos H. Papadimitriou, and Karthik Narasimhan · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Closest in time.
End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, and Tom Goldstein · 2022
Closest in time.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Closest in time.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A Ortega · 2022
Closest in time.
A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level
Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, Roman Wang, Nikhil Singh, Taylor L. Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, and Gilbert Strang · 2022
Closest in time.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Closest in time.
Transformer language models without positional encodings still learn positional information
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy · 2022
Closest in time.
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur · 2022
Closest in time.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Closest in time.
Transformers are sample efficient world models
Vincent Micheli, Eloi Alonso, and François Fleuret · 2022
Closest in time.
A mechanistic interpretability analysis of grokking
Neel Nanda and Tom Lieberum · 2022
Closest in time.
Eshaan Nichani, Yu Bai, and Jason D Lee · 2022
Closest in time.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.
A survey on non-autoregressive generation for neural machine translation and beyond
Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, and Tie-yan Liu · 2022
Closest in time.
Unveiling Transformers with LEGO: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Closest in time.