Fetching the paper…
Reading the bibliography…
Algorithmic generalization in machine learning refers to the ability to learn the underlying algorithm that generates data in a way that generalizes out-of-distribution.
2016
Earlier work this paper cites.
L. Kaiser and I. Sutskever, “Neural gpus learn algorithms,” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems
2017
Earlier work this paper cites.
G. Weiss, Y. Goldberg, and E. Yahav, “On the practical computational power of finite precision rnns for language recognition,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
2018
Cited alongside, same era.
F. Chollet, “On the measure of intelligence,” CoRR
2019
Cited alongside, same era.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser, “Universal transformers,” in Proceedings of International Conference on Learning Representations
2019
Cited alongside, same era.
M. Suzgun, S. Gehrmann, Y. Belinkov12, and S. M. Shieber, “Lstm networks can perform dynamic counting,” ACL 2019
2019
Cited alongside, same era.
M. Hahn, “Theoretical limitations of self-attention in neural sequence models,” Transactions of the Association for Computational Linguistics
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Cognolato and A. Testolin, “Transformers discover an elementary calculation system exploiting local attention and grid-like problem representation,” in 2022 International Joint Conference on Neural Networks (IJCNN)
2022
Later among the works it cites.
N. El-Naggar, P. Madhyastha, and T. Weyde, “Exploring the long-term generalization of counting behavior in rnns,” in I Can’t Believe It’s Not Better Workshop: Understanding Deep Learning Through Empirical Falsification
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…