Fetching the paper…
Reading the bibliography…
Inspired from human cognition, machine learning systems are gradually revealing advantages of sparser and more modular architectures.
The hungarian method for the assignment problem
Kuhn, H. W · 1955
Earlier work this paper cites.
In the theatre of consciousness. global workspace theory, a rigorous scientific theory of consciousness
Baars, B. J · 1997
Earlier work this paper cites.
Twenty years of mixture of experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D · 2012
Earlier work this paper cites.
Deep learning of representations: Looking forward
Bengio, Y · 2013
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2016
Earlier work this paper cites.
Bengio, Y · 2017
Earlier work this paper cites.
What is consciousness, and could machines have it?
Dehaene, S., Lau, H., and Kouider, S · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Hu, R., Andreas, J., Rohrbach, M., Darrell, T., and Saenko, K · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., et al · 2018
Cited alongside, same era.
Disentangling by factorising
Kim, H. and Mnih, A · 2018
Cited alongside, same era.
Neural relational inference for interacting systems
Kipf, T., Fetaya, E., Wang, K.-C., Welling, M., and Zemel, R · 2018
Cited alongside, same era.
Re-examining routing networks for multi-task learning
Cui, L. and Jaech, A · 2020
Later among the works it cites.
Inductive biases for deep learning of higher-level cognition
Goyal, A. and Bengio, Y · 2020
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Learning to combine top-down and bottom-up signals in recurrent neural networks with attention over modules
Mittal, S., Lamb, A., Goyal, A., Voleti, V., Shanahan, M., Lajoie, G., Mozer, M., and Bengio, Y · 2020
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Santoro, A., Faulkner, R., Raposo, D., Rae, J., Chrzanowski, M., Weber, T., Wierstra, D., Vinyals, O., Pascanu, R., and Lillicrap, T · 2018
Cited alongside, same era.
A meta-transfer objective for learning to disentangle causal mechanisms
Bengio, Y., Deleu, T., Rahaman, N., Ke, N. R., Lachapelle, S., Bilaniuk, O., Goyal, A., and Pal, C · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B · 2019
Cited alongside, same era.
Flexible multi-task networks by learning parameter allocation
Maziarz, K., Kokiopoulou, E., Gesmundo, A., Sbaiz, L., Bartok, G., and Berent, J · 2019
Cited alongside, same era.
Routing networks and the challenges of modular and compositional computation
Rosenbaum, C., Cases, I., Riemer, M., and Klinger, T · 2019
Cited alongside, same era.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Csordás, R., van Steenkiste, S., and Schmidhuber, J · 2020
Cited alongside, same era.
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Goyal, A., Didolkar, A., Ke, N. R., Blundell, C., Beaudoin, P., Heess, N., Mozer, M., and Bengio, Y · 2021
Later among the works it cites.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Ke, N. R., Didolkar, A. R., Mittal, S., Goyal, A., Lajoie, G., Bauer, S., Rezende, D. J., Mozer, M. C., Bengio, Y., and Pal, C · 2021
Later among the works it cites.
Fast and slow learning of recurrent independent mechanisms
Madan, K., Ke, R. N., Goyal, A., Schölkopf, B. B., and Bengio, Y · 2021
Later among the works it cites.
Compositional attention: Disentangling search and retrieval
Mittal, S., Raparthy, S. C., Rish, I., Bengio, Y., and Lajoie, G · 2021
Later among the works it cites.
Dynamic inference with neural interpreters
Rahaman, N., Gondal, M. W., Joshi, S., Gehler, P., Bengio, Y., Locatello, F., and Schölkopf, B · 2021
Later among the works it cites.