Fetching the paper…
Reading the bibliography…
Transformers have impressive generalization capabilities on tasks with a fixed context length.
Sequential neural networks as automata
William Merrill. 2019 · 1906
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky. 1956 · 1956
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Earlier work this paper cites.
A survey of neural networks and formal languages
Joshua Ackerman and George Cybenko. 2020 · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden M. Lake and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Earlier work this paper cites.
How can self-attention networks recognize dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang. 2020 · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet. 2020 · 2020
Cited alongside, same era.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Cited alongside, same era.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2022 · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Later among the works it cites.
A generalist neural algorithmic learner
Borja Ibarz, Vitaly Kurin, George Papamakarios, Kyriacos Nikiforou, Mehdi Bennani, Róbert Csordás, Andrew Joseph Dudzik, Matko Bosnjak, Alex Vitvitskyi, Yulia Rubanova, Andreea Deac, Beatrice Bevilacqua, Yaroslav Ganin, Charles Blundell, and Petar Velickovic. 2022 · 2022
Later among the works it cites.
Systematic generalization and emergent structures in transformers trained on structured tasks
Yuxuan Li and James L. McClelland. 2022 · 2022
Later among the works it cites.
Log-precision transformers are constant-depth uniform threshold circuits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021 · 2021
Cited alongside, same era.
Random features strengthen graph neural networks
Ryoma Sato, Makoto Yamada, and Hisashi Kashima. 2021 · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021 · 2021
Cited alongside, same era.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2021 · 2021
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Cited alongside, same era.
Compositional generalization by learning analytical expressions
Qian Liu, Shengnan An, Jian-Guang Lou, Bei Chen, Zeqi Lin, Yan Gao, Bin Zhou, Nanning Zheng, and Dongmei Zhang. 2020a
Cited in the paper.
William Merrill and Ashish Sabharwal. 2022 · 2022
Later among the works it cites.
Making transformers solve compositional tasks
Santiago Ontañón, Joshua Ainslie, Zachary Fisher, and Vaclav Cvicek. 2022 · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis. 2022 · 2022
Later among the works it cites.
A generalist agent
Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. 2022 · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega. 2023 · 2023
Closest in time.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2023
Closest in time.