Fetching the paper…
Reading the bibliography…
Deep learning has seen a movement away from representing examples with a monolithic hidden state towards a richly structured state.
CATER: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 1910
Earlier work this paper cites.
The modularity of mind
Jerry A Fodor · 1983
Earlier work this paper cites.
Vehicles: Experiments in synthetic psychology
Valentino Braitenberg · 1986
Earlier work this paper cites.
Society of mind
Marvin Minsky · 1988
Earlier work this paper cites.
A framework for the cooperation of learning algorithms
Léon Bottou and Patrick Gallinari · 1991
Earlier work this paper cites.
Intelligence without representation
Rodney A Brooks · 1991
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
A cognitive theory of consciousness
Bernard J Baars · 1993
Earlier work this paper cites.
In the theatre of consciousness. global workspace theory, a rigorous scientific theory of consciousness
Bernard J Baars · 1997
Earlier work this paper cites.
Modular neural networks and self-decomposition
Eric Ronco, Henrik Gollee, and Peter J Gawthrop · 1997
Earlier work this paper cites.
A neuronal model of a global workspace in effortful cognitive tasks
Stanislas Dehaene, Michel Kerszberg, and Jean-Pierre Changeux · 1998
Earlier work this paper cites.
Theories of access consciousness
Michael D. Colagrosso and Michael C Mozer · 2005
Earlier work this paper cites.
Applying global workspace theory to the frame problem
Murray Shanahan and Bernard Baars · 2005
Earlier work this paper cites.
A cognitive architecture that combines internal simulation with a global workspace
Murray Shanahan · 2006
Earlier work this paper cites.
Equilateral triangles: A challenge for connectionist vision
S. Ahmad and S. Omohundro · 2009
Earlier work this paper cites.
Embodiment and the inner life: Cognition and Consciousness in the Space of Possible Minds
Murray Shanahan · 2010
Earlier work this paper cites.
Experimental and theoretical approaches to conscious processing
Stanislas Dehaene and Jean-Pierre Changeux · 2011
Earlier work this paper cites.
The brain’s connective core and its role in animal cognition
Murray Shanahan · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Neural programmer-interpreters
Scott Reed and Nando De Freitas · 2015
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What is consciousness, and could machines have it?
S. Dehaene, H. Lau, and S. Kouider · 2017
Cited alongside, same era.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Cited alongside, same era.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Cited alongside, same era.
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2019
Later among the works it cites.
Routing networks and the challenges of modular and compositional computation
Clemens Rosenbaum, Ignacio Cases, Matthew Riemer, and Tim Klinger · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Mikhail S Burtsev and Grigory V Sapunov · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Inductive biases for deep learning of higher-level cognition
Anirudh Goyal and Yoshua Bengio · 2020
Later among the works it cites.
Object files and schemata: Factorizing declarative and procedural knowledge in dynamical systems
Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin, Sergey Levine, Charles Blundell, Yoshua Bengio, and Michael Mozer · 2020
Later among the works it cites.
karpathy/mingpt, Aug 2020
Andrej Karpathy · 2020
Later among the works it cites.
S2rms: Spatially structured recurrent modules
Nasim Rahaman, Anirudh Goyal, Muhammad Waleed Gondal, Manuel Wuthrich, Stefan Bauer, Yash Sharma, Yoshua Bengio, and Bernhard Schölkopf · 2020
Later among the works it cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.
Transformers with competitive ensembles of independent mechanisms, 2021
Alex Lamb, Di He, Anirudh Goyal, Guolin Ke, Chien-Feng Liao, Mirco Ravanelli, and Yoshua Bengio · 2021
Closest in time.
Meta attention networks: Meta-learning attention to modulate information between recurrent independent mechanisms
Kanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf, and Yoshua Bengio · 2021
Closest in time.