Fetching the paper…
Reading the bibliography…
This paper investigates the failure cases and out-of-distribution behavior of transformers trained on matrix inversion and eigenvalue decomposition.
Random Matrices
Madan Lal Mehta · 2004
Earlier work this paper cites.
Matrix Computations
Gene H. Golub and Charles F. van Loan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2019
Cited alongside, same era.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2020
Cited alongside, same era.
It’s not what machines can learn, it’s what we cannot teach
Gal Yehuda, Moshe Gabel, and Assaf Schuster · 2020
Cited alongside, same era.
Neural symbolic regression that scales
Luca Biggio, Tommaso Bendinelli, Alexander Neitz, Aurelien Lucchi, and Giambattista Parascandolo · 2021
Cited alongside, same era.
Linear algebra with transformers
François Charton · 2021
Later among the works it cites.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Later among the works it cites.
Simplifying polylogarithms with machine learning
Aurélien Dersy, Matthew D. Schwartz, and Xiaoyuan Zhang · 2022
Closest in time.
Symbolic brittleness in sequence models: on systematic generalization in symbolic mathematics
Sean Welleck, Peter West, Jize Cao, and Yejin Choi · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…