Fetching the paper…
Reading the bibliography…
Can transformers generalize efficiently on problems that require dealing with examples with different levels of difficulty? We introduce a new task tailored to assess generalization over different complexities and present results that indicate that standard transformers face challenges in solving these tasks.
Compositionality
Zoltán Gendler Szabó · 2008
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Hypernetworks
David Ha, Andrew M. Dai, and Quoc V. Le · 2017
Earlier work this paper cites.
Recurrent scene parsing with perspective understanding in the loop
Shu Kong and Charless C. Fowlkes · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Earlier work this paper cites.
Shallow-Deep Networks: Understanding and mitigating network overthinking
Yiğitcan Kaya, Sanghyun Hong, and Tudor Dumitras · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Earlier work this paper cites.
Resolution adaptive networks for efficient inference
Le Yang, Yizeng Han, Xi Chen, Shiji Song, Jifeng Dai, and Gao Huang · 2020
Earlier work this paper cites.
Transferring inductive biases through knowledge distillation, 2021
Samira Abnar, Mostafa Dehghani, and Willem H. Zuidema · 2021
Earlier work this paper cites.
Deep learning through the lens of example difficulty
Robert Baldock, Hartmut Maennel, and Behnam Neyshabur · 2021
Earlier work this paper cites.
Pondernet: Learning to ponder
Andrea Banino, Jan Balaguer, and Charles Blundell · 2021
Cited alongside, same era.
Deep learning for ai
Yoshua Bengio, Yann Lecun, and Geoffrey Hinton · 2021
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
GLaM: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten P Bosma, Zongwei Zhou, Tao Wang, Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc Le, Yonghui Wu, Zhifeng Chen, and Claire Cui · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Later among the works it cites.
Hypermixer: An mlp-based low cost alternative to transformers
Florian Mai, Arnaud Pannatier, Fabio Fehr, Haolin Chen, François Marelli, François Fleuret, and James Henderson · 2022
Later among the works it cites.
Understanding the robustness of multi-exit models under common corruptions
Akshay Mehra, Skyler Seto, Navdeep Jaitly, and Barry-John Theobald · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse is enough in scaling transformers
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, LUKASZ KAISER, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Cited alongside, same era.
Adaptive inference through early-exit networks: Design, challenges and directions
Stefanos Laskaridis, Alexandros Kouris, and Nicholas D. Lane · 2021
Cited alongside, same era.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Cited alongside, same era.
Chiyuan Zhang, Maithra Raghu, Jon M. Kleinberg, and Samy Bengio · 2021
Cited alongside, same era.
Estimating example difficulty using variance of gradients
Chirag Agarwal, Daniel D’souza, and Sara Hooker · 2022
Cited alongside, same era.
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Later among the works it cites.
Generalization on the unseen, logic reasoning and degree curriculum
Emmanuel Abbe, Samy Bengio, Aryo Lotfi, and Kevin Rizk · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Length generalization in arithmetic transformers, 2023
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton · 2023
Closest in time.
Recursion of thought: Divide and conquer reasoning with language models, 2023
Soochan Lee and Gunhee Kim · 2023
Closest in time.
Jonas Pfeiffer, Sebastian Ruder, Ivan Vuli’c, and Edoardo M. Ponti · 2023
Closest in time.
Learning multi-step reasoning by solving arithmetic tasks
Tianduo Wang and Wei Lu · 2023
Closest in time.
Adaptive computation with elastic input sequence, 2023
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Yi Tay, Mostafa Dehghani, and Yang You · 2023
Closest in time.