Fetching the paper…
Reading the bibliography…
Current deep learning approaches have shown good in-distribution generalization performance, but struggle with out-of-distribution generalization.
Raven’s progressive matrices
Raven, J. C. and Court, J · 1938
Earlier work this paper cites.
Cognitive maps in rats and men
Tolman, E · 1948
Earlier work this paper cites.
Relations relating relations
Goldstone, R. L., Gentner, D., and Medin, D. L · 1989
Earlier work this paper cites.
What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test
Carpenter, P. A., Just, M. A., and Shell, P · 1990
Earlier work this paper cites.
Separate visual pathways for perception and action
Goodale, M. A. and Milner, A. D · 1992
Earlier work this paper cites.
Memory, amnesia, and the hippocampal system
Cohen, N. and Eichenbaum, H · 1993
Earlier work this paper cites.
Rule learning by seven-month-old infants
Marcus, G. F., Vijayan, S., Rao, S. B., and Vishton, P. M · 1999
Earlier work this paper cites.
Grounding conceptual knowledge in modality-specific systems
Barsalou, L. W., Kyle Simmons, W., Barbey, A. K., and Wilson, C. D · 2003
Earlier work this paper cites.
Beyond common features: The role of roles in determining similarity
Jones, M. and Love, B. C · 2006
Earlier work this paper cites.
Learning representations that support extrapolation
Webb, T. W., Dulberg, Z., Frankland, S. M., Petrov, A. A., O’Reilly, R. C., and Cohen, J. D · 2007
Earlier work this paper cites.
The graph neural network model
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G · 2008
Earlier work this paper cites.
Dorsal and ventral streams across sensory modalities
Sedda, A. and Scarpina, F · 2012
Earlier work this paper cites.
Emergent symbols through binding in external memory
Webb, T. W., Sinha, I., and Cohen, J. D · 2012
Earlier work this paper cites.
Indirection and symbol-like processing in the prefrontal cortex and basal ganglia
Kriete, T., Noelle, D. C., Cohen, J. D., and O’Reilly, R. C · 2013
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Cited alongside, same era.
Pritzel, A., Uria, B., Srinivasan, S., Puigdomenech, A., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T · 2017
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Generalization of reinforcement learners with working and episodic memory
Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., and Blundell, C · 2019
Later among the works it cites.
Wav2letter++: A fast open-source speech recognition system
Pratap, V., Hannun, A., Xu, Q., Cai, J., Kahn, J., Synnaeve, G., Liptchinsky, V., and Collobert, R · 2019
Later among the works it cites.
An explicitly relational neural network architecture
Shanahan, M., Nikiforou, K., Creswell, A., Kaplanis, C., Barrett, D., and Garnelo, M · 2019
Later among the works it cites.
The tolman-eichenbaum machine: Unifying space and relational memory through generalisation in the hippocampal formation
Whittington, J. C., Muller, T. H., Mark, S., Chen, G., Barry, C., Burgess, N., and Behrens, T. E · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y · 2017
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., et al · 2018
Cited alongside, same era.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D · 2018
Cited alongside, same era.
Not-so-clevr: learning same–different relations strains feedforward neural networks
Kim J, Ricci M, S. T · 2018
Cited alongside, same era.
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2019
Later among the works it cites.
Raven: A dataset for relational and analogical visual reasoning
Zhang, C., Gao, F., Jia, B., Zhu, Y., and Zhu, S.-C · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Grounded language learning fast and slow
Hill, F., Tieleman, O., von Glehn, T., Wong, N., Merzic, H., and Clark, S · 2020
Later among the works it cites.
Cogs: A compositional generalization challenge based on semantic interpretation
Kim, N. and Linzen, T · 2020
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Learning to combine top-down and bottom-up signals in recurrent neural networks with attention over modules
Mittal, S., Lamb, A., Goyal, A., Voleti, V., Shanahan, M., Lajoie, G., Mozer, M., and Bengio, Y · 2020
Later among the works it cites.
The eos decision and length extrapolation
Newman, B., Hewitt, J., Liang, P., and Manning, C. D · 2020
Later among the works it cites.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Ke, N. R., Didolkar, A. R., Mittal, S., Goyal, A., Lajoie, G., Bauer, S., Rezende, D. J., Mozer, M. C., Bengio, Y., and Pal, C · 2021
Later among the works it cites.
Compositional attention: Disentangling search and retrieval
Mittal, S., Raparthy, S. C., Rish, I., Bengio, Y., and Lajoie, G · 2021
Later among the works it cites.
Investigating the limitations of transformers with simple arithmetic tasks
Nogueira, R., Jiang, Z., and Lin, J · 2021
Later among the works it cites.