Fetching the paper…
Reading the bibliography…
We examine how transformers cope with two challenges: learning basic integer arithmetic, and generalizing to longer sequences than seen during training.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov · 2015
Earlier work this paper cites.
Łukasz Kaiser and Ilya Sutskever · 2015
Earlier work this paper cites.
Dex: Deep expectation of apparent age from a single image
Rasmus Rothe, Radu Timofte, and Luc Van Gool · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Investigating the ability of neural networks to learn simple modular arithmetic
Theodoros Palamas · 2017
Earlier work this paper cites.
Lcr-net: Localization-classification-regression for human pose
Gregory Rogez, Philippe Weinzaepfel, and Cordelia Schmid · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
An improved relative self-attention mechanism for transformer with application to music generation
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, and Douglas Eck · 2018
Earlier work this paper cites.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Andrew Trask, Felix Hill, Scott Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Location attention for extrapolation to longer sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, and Elia Bruni · 2019
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Cited alongside, same era.
Solving math word problems with double-decoder transformer
Yuanliang Meng and Anna Rumshisky · 2019
Cited alongside, same era.
Analysis of positional encodings for neural machine translation
Jan Rosendahl, Viet Anh Khoa Tran, Weiyue Wang, and Hermann Ney · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Shape: Shifted absolute position embedding for transformers
Shun Kiyono, Sosuke Kobayashi, Jun Suzuki, and Kentaro Inui · 2021
Later among the works it cites.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Teaching temporal logics to neural networks
Christopher Hahn, Frederik Schmitt, Jens U Kreber, Markus N Rabe, and Bernd Finkbeiner · 2020
Cited alongside, same era.
Improve transformer models with better relative position embeddings
Zhiheng Huang, Davis Liang, Peng Xu, and Bing Xiang · 2020
Cited alongside, same era.
The eos decision and length extrapolation
Benjamin Newman, John Hewitt, Percy Liang, and Christopher D Manning · 2020
Cited alongside, same era.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Cited alongside, same era.
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Later among the works it cites.
End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking, 2022
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, and Tom Goldstein · 2022
Later among the works it cites.
Mirelle Bueno, Carlos Gemmel, Jeffrey Dalton, Roberto Lotufo, and Rodrigo Nogueira · 2022
Later among the works it cites.
Simplifying polylogarithms with machine learning, 2022
Aurélien Dersy, Matthew D. Schwartz, and Xiaoyuan Zhang · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue, 2022
J Schulman, B Zoph, C Kim, J Hilton, J Menick, J Weng, JFC Uribe, L Fedus, L Metz, M Pokorny, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Salsa: Attacking lattice cryptography with transformers
Emily Wenger, Mingjie Chen, François Charton, and Kristin Lauter · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Later among the works it cites.
Teaching algorithmic reasoning via in-context learning
Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron Courville, Behnam Neyshabur, and Hanie Sedghi · 2022
Later among the works it cites.