Fetching the paper…
Reading the bibliography…
Modern neural network architectures can leverage large amounts of data to generalize well within the training distribution.
Using anytime algorithms in intelligent systems
Shlomo Zilberstein · 1996
Earlier work this paper cites.
Raven’s progressive matrices and vocabulary scales
John C Raven and John Hugh Court · 1998
Earlier work this paper cites.
An introduction to knowledge engineering
Simon L Kendal and Malcolm Creen · 2007
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
On causal and anticausal learning
Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris Mooij · 2012
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Earlier work this paper cites.
Measuring the tendency of cnns to learn surface statistical regularities, 2017
Jason Jo and Yoshua Bengio · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Measuring abstract reasoning in neural networks, 2018
David G. T. Barrett, Felix Hill, Adam Santoro, Ari S. Morcos, and Timothy Lillicrap · 2018
Earlier work this paper cites.
Automatically composing representation transformations as a means for generalization
Michael B Chang, Abhishek Gupta, Sergey Levine, and Thomas L Griffiths · 2018
Earlier work this paper cites.
Deep learning for classical japanese literature, 2018
Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization, 2018
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Cited alongside, same era.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Cited alongside, same era.
Rearranging the familiar: Testing compositional generalization in recurrent networks
Joao Loula, Marco Baroni, and Brenden M Lake · 2018
Cited alongside, same era.
Learning independent causal mechanisms
G. Parascandolo, N. Kilbertus, M. Rojas-Carulla, and B. Schölkopf · 2018
Cited alongside, same era.
Improving generalization for abstract reasoning tasks using disentangled feature representations
Xander Steenbrugge, Sam Leroux, Tim Verbelen, and Bart Dhoedt · 2018
Repulsive attention: Rethinking multi-head attention as bayesian inference
Bang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, and Changyou Chen · 2020
Later among the works it cites.
Image generators with conditionally-independent pixel synthesis, 2020
Ivan Anokhin, Kirill Demochkin, Taras Khakhulin, Gleb Sterkin, Victor Lempitsky, and Denis Korzhenkov · 2020
Later among the works it cites.
A theory of independent mechanisms for extrapolation in generative models
Michel Besserve, Rémy Sun, Dominik Janzing, and Bernhard Schölkopf · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Sjoerd Van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Systematic generalization: What is required and can it be learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville · 2019
Cited alongside, same era.
A meta-transfer objective for learning to disentangle causal mechanisms, 2019
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher Pal · 2019
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2019
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space, 2019
Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le · 2019
Cited alongside, same era.
Universal transformers, 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Anirudh Goyal and Yoshua Bengio · 2020
Later among the works it cites.
Object files and schemata: Factorizing declarative and procedural knowledge in dynamical systems
Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin, Sergey Levine, Charles Blundell, Yoshua Bengio, and Michael Mozer · 2020
Later among the works it cites.
Gaussian error linear units (gelus), 2020
Dan Hendrycks and Kevin Gimpel · 2020
Later among the works it cites.
Object-centric learning with slot attention, 2020
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Few-shot sequence learning with transformers
Lajanugen Logeswaran, Ann Lee, Myle Ott, Honglak Lee, Marc’Aurelio Ranzato, and Arthur Szlam · 2020
Later among the works it cites.
S2rms: Spatially structured recurrent modules, 2020
Nasim Rahaman, Anirudh Goyal, Muhammad Waleed Gondal, Manuel Wuthrich, Stefan Bauer, Yash Sharma, Yoshua Bengio, and Bernhard Schölkopf · 2020
Later among the works it cites.
Abstract diagrammatic reasoning with multiplex graph networks
Duo Wang, Mateja Jamnik, and Pietro Lio · 2020
Later among the works it cites.
Yuhuai Wu, Honghua Dong, Roger Grosse, and Jimmy Ba · 2020
Later among the works it cites.
Neural event semantics for grounded language understanding
Shyamal Buch, Li Fei-Fei, and Noah D. Goodman · 2021
Closest in time.
An empirical study of training self-supervised vision transformers, 2021
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Anirudh Goyal, Aniket Didolkar, Nan Rosemary Ke, Charles Blundell, Philippe Beaudoin, Nicolas Heess, Michael Mozer, and Yoshua Bengio · 2021
Closest in time.
Perceiver: General perception with iterative attention, 2021
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.
Transformers with competitive ensembles of independent mechanisms
Alex Lamb, Di He, Anirudh Goyal, Guolin Ke, Chien-Feng Liao, Mirco Ravanelli, and Yoshua Bengio · 2021
Closest in time.
Towards causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, R. Nan Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Closest in time.
Training data-efficient image transformers and distillation through attention, 2021
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Closest in time.
Going deeper with image transformers, 2021
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Closest in time.