Fetching the paper…
Reading the bibliography…
The predictions of small transformers, trained to calculate the greatest common divisor (GCD) of two positive integers, can be fully characterized by looking at model inputs and outputs.
Optimal depth neural networks for multiplication and related problems
Kai-Yeung Siu and Vwani Roychowdhury · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Łukasz Kaiser and Ilya Sutskever · 2015
Earlier work this paper cites.
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves · 2015
Earlier work this paper cites.
Learning simple algorithms from examples, 2015
Wojciech Zaremba, Tomas Mikolov, Armand Joulin, and Rob Fergus · 2015
Earlier work this paper cites.
Investigating the ability of neural networks to learn simple modular arithmetic
Theodoros Palamas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Andrew Trask, Felix Hill, Scott Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2019
Earlier work this paper cites.
Solving math word problems with double-decoder transformer
Yuanliang Meng and Anna Rumshisky · 2019
Cited alongside, same era.
Analysing mathematical reasoning abilities of neural models, 2019
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2019
Cited alongside, same era.
Learning advanced mathematical computations from examples
François Charton, Amaury Hayat, and Guillaume Lample · 2020
Cited alongside, same era.
Linear algebra with transformers
François Charton · 2021
Cited alongside, same era.
Solving arithmetic word problems with transformers and preprocessing of problem text
Kaden Griffith and Jugal Kalita · 2021
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Teaching algorithmic reasoning via in-context learning
Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron Courville, Behnam Neyshabur, and Hanie Sedghi · 2022
Later among the works it cites.
Mathematics, word problems, common sense, and artificial intelligence, 2023
Ernest Davis · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality, 2023
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi · 2023
Closest in time.
Grokking modular arithmetic, 2023
Andrey Gromov · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Cited alongside, same era.
Transformer-based Machine Learning for Fast SAT Solvers and Logic Synthesis
Feng Shi, Chonghan Lee, Mohammad Khairul Bashar, Nikhil Shukla, Song-Chun Zhu, and Vijaykrishnan Narayanan · 2021
Cited alongside, same era.
Symbolic Brittleness in Sequence Models: on Systematic Generalization in Symbolic Mathematics, 2021
Sean Welleck, Peter West, Jize Cao, and Yejin Choi · 2021
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning, 2022
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud, Max Tegmark, and Mike Williams · 2022
Cited alongside, same era.
Question 75 (solution)
Ernesto Cesàro
Cited in the paper.
What is my math transformer doing? – three results on interpretability and generalization
François Charton
Cited in the paper.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
An investigation into neural arithmetic logic modules
Bhumika Mistry · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks, 2023
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Closest in time.