Fetching the paper…
Reading the bibliography…
Multi-head, key-value attention is the backbone of the widely successful Transformer model and its variants.
In the theatre of consciousness. global workspace theory, a rigorous scientific theory of consciousness
Bernard J Baars · 1997
Earlier work this paper cites.
Equilateral triangles: A challenge for connectionist vision
S. Ahmad and S. Omohundro · 2009
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Yoshua Bengio · 2017
Earlier work this paper cites.
What is consciousness, and could machines have it?
S. Dehaene, H. Lau, and S. Kouider · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David GT Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Earlier work this paper cites.
Memorize or generalize? searching for a compositional RNN in a haystack
Adam Liska, Germán Kruszewski, and Marco Baroni · 2018
Earlier work this paper cites.
Compositional attention networks for interpretability in natural language question answering
Muru Selvakumar, Suriyadeepan Ramamoorthy, Vaidheeswaran Archana, and Malaikannan Sankarasubbu · 2018
Earlier work this paper cites.
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Cited alongside, same era.
A meta-transfer objective for learning to disentangle causal mechanisms
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher Pal · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2019
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Object files and schemata: Factorizing declarative and procedural knowledge in dynamical systems
Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin, Sergey Levine, Charles Blundell, Yoshua Bengio, and Michael Mozer · 2020
Later among the works it cites.
Compositionality decomposed: how do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Later among the works it cites.
Untangling tradeoffs between recurrence and self-attention in artificial neural networks
Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal ALIAS PARTH GOYAL, Kyle Goyette, Yoshua Bengio, and Guillaume Lajoie · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, et al · 2019
Cited alongside, same era.
Compositional generalization for primitive substitutions
Yuanpeng Li, Liang Zhao, Jianyu Wang, and Joel Hestness · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Cited alongside, same era.
Compositional generalization via neural-symbolic stack machines
Xinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song, and Denny Zhou · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
David Ding, Felix Hill, Adam Santoro, and Matt Botvinick · 2020
Cited alongside, same era.
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Learning to combine top-down and bottom-up signals in recurrent neural networks with attention over modules
Sarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti, Murray Shanahan, Guillaume Lajoie, Michael Mozer, and Yoshua Bengio · 2020
Later among the works it cites.
The eos decision and length extrapolation
Benjamin Newman, John Hewitt, Percy Liang, and Christopher D Manning · 2020
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Emergent symbols through binding in external memory
Taylor W Webb, Ishan Sinha, and Jonathan D Cohen · 2020
Later among the works it cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Closest in time.
Systematic evaluation of causal discovery in visual model based reinforcement learning, 2021
Nan Rosemary Ke, Aniket Rajiv Didolkar, Sarthak Mittal, Anirudh Goyal, Guillaume Lajoie, Stefan Bauer, Danilo Jimenez Rezende, Michael Curtis Mozer, Yoshua Bengio, and Christopher Pal · 2021
Closest in time.
Transformers with competitive ensembles of independent mechanisms
Alex Lamb, Di He, Anirudh Goyal, Guolin Ke, Chien-Feng Liao, Mirco Ravanelli, and Yoshua Bengio · 2021
Closest in time.
Fast and slow learning of recurrent independent mechanisms
Kanika Madan, Rosemary Nan Ke, Anirudh Goyal, Bernhard Bernhard Schölkopf, and Yoshua Bengio · 2021
Closest in time.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Closest in time.