Fetching the paper…
Reading the bibliography…
Compositional generalization, the ability of an agent to generalize to unseen combinations of latent factors, is easy for humans but hard for deep neural networks.
“Invariant risk minimization”, 2019
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani and David Lopez-Paz · 1907
Earlier work this paper cites.
“The origin of speech”
Charles Hockett · 1960
Earlier work this paper cites.
“Connectionism and cognitive architecture: A critical analysis”
Jerry Fodor and Zenon Pylyshyn · 1988
Earlier work this paper cites.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning”
Ronald Williams · 1992
Earlier work this paper cites.
“Long short-term memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“The information bottleneck method”
Naftali Tishby, Fernando Pereira and William Bialek · 1999
Earlier work this paper cites.
“The compositionality papers”
Jerry Fodor and Ernest Lepore · 2002
Earlier work this paper cites.
“Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond”
Bernhard Schölkopf and Alexander Smola · 2002
Earlier work this paper cites.
“Measuring Statistical Dependence with Hilbert-Schmidt Norms”
Arthur Gretton, Olivier Bousquet, Alex Smola and Bernhard Schölkopf · 2005
Earlier work this paper cites.
“Understanding linguistic evolution by visualizing the emergence of topographic mappings”
Henry Brighton and Simon Kirby · 2006
Earlier work this paper cites.
“Measuring and testing dependence by correlation of distances”
Gábor. Székely, Maria. Rizzo and Nail. Bakirov · 2007
Earlier work this paper cites.
“Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language”
Simon Kirby, Hannah Cornish and Kenny Smith · 2008
Earlier work this paper cites.
“Extracting and composing robust features with denoising autoencoders”
Pascal Vincent, Hugo Larochelle, Yoshua Bengio and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
“Iterated learning and the cultural ratchet”
Aaron Beppu and Thomas Griffiths · 2009
Earlier work this paper cites.
“A generalization of transformer networks to graphs”
Vijay Dwivedi and Xavier Bresson · 2012
Earlier work this paper cites.
“Estimating or propagating gradients through stochastic neurons for conditional computation”, 2013
Yoshua Bengio, Nicholas Léonard and Aaron Courville · 2013
Earlier work this paper cites.
“Equivalence of distance-based and RKHS-based statistics in hypothesis testing”
Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton and Kenji Fukumizu · 2013
Earlier work this paper cites.
“Distilling the knowledge in a neural network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Compression and communication in the cultural evolution of linguistic structure”
Simon Kirby, Monica Tamariz, Hannah Cornish and Kenny Smith · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander. Berg and Li Fei-Fei · 2015
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Categorical reparameterization with Gumbel-softmax”, 2016
Eric Jang, Shixiang Gu and Ben Poole · 2016
Earlier work this paper cites.
“Rdkit: Open-source cheminformatics software”, 2016
Greg Landrum · 2016
Earlier work this paper cites.
“Partial differential equations in action: from modeling to theory”
Sandro Salsa · 2016
Earlier work this paper cites.
“Mastering the game of Go with deep neural networks and tree search”
David Silver, Aja Huang, Chris Maddison, Arthur Guez, Laurent Sifre, George Van, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel and Demis Hassabis · 2016
Earlier work this paper cites.
“Neural message passing for quantum chemistry”
Justin Gilmer, Samuel Schoenholz, Patrick Riley, Oriol Vinyals and George Dahl · 2017
Earlier work this paper cites.
“Inductive representation learning on large graphs”
Will Hamilton, Zhitao Ying and Jure Leskovec · 2017
Earlier work this paper cites.
“Emergence of language with multi-agent games: Learning to communicate with sequences of symbols”
Serhii Havrylov and Ivan Titov · 2017
Earlier work this paper cites.
“Semi-Supervised Classification with Graph Convolutional Networks”
Thomas. Kipf and Max Welling · 2017
Cited alongside, same era.
“dSprites: Disentanglement testing Sprites dataset”, 2017
Loic Matthey, Irina Higgins, Demis Hassabis and Alexander Lerchner · 2017
Cited alongside, same era.
“Opening the black box of deep neural networks via information”, 2017
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
“Neural discrete representation learning”
Aaron van Oord, Oriol Vinyals and Koray Kavukcuoglu · 2017
Cited alongside, same era.
“3D Shapes Dataset”, 2018
Chris Burgess and Hyunjik Kim · 2018
Cited alongside, same era.
“Born again neural networks”
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti and Anima Anandkumar · 2018
Cited alongside, same era.
“Discrete-valued neural communication”
Dianbo Liu, Alex Lamb, Kenji Kawaguchi, Anirudh Goyal, Chen Sun, Michael Mozer and Yoshua Bengio · 2021
Later among the works it cites.
“Iterated learning for emergent systematicity in VQA”
Ankit Vani, Max Schwarzer, Yuchen Lu, Eeshan Dhekane and Aaron Courville · 2021
Later among the works it cites.
“BEiT: BERT Pre-Training of Image Transformers”
Hangbo Bao, Li Dong, Songhao Piao and Furu Wei · 2022
Later among the works it cites.
“Long Range Graph Benchmark”
Vijay Dwivedi, Ladislav Rampášek, Mikhail Galkin, Ali Parviz, Guy Wolf, Anh Luu and Dominique Beaini · 2022
Later among the works it cites.
“Simple GNN Regularisation for 3D Molecular Property Prediction and Beyond”
Jonathan Godwin, Michael Schaarschmidt, Alexander Gaunt, Alvaro Sanchez-Gonzalez, Yulia Rubanova, Petar Veličković, James Kirkpatrick and Peter Battaglia · 2022
Later among the works it cites.
“Inductive biases for deep learning of higher-level cognition”
Anirudh Goyal and Yoshua Bengio · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Towards a definition of disentangled representations”, 2018
Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende and Alexander Lerchner · 2018
Cited alongside, same era.
“The Book of Why: The New Science of Cause and Effect”
Judea Pearl and Dana Mackenzie · 2018
Cited alongside, same era.
“MoleculeNet: a benchmark for molecular machine learning”
Zhenqin Wu, Bharath Ramsundar, Evan Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh Pappu, Karl Leswing and Vijay Pande · 2018
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“On the Transfer of Inductive Bias from Simulation to the Real World: a New Disentanglement Dataset”
Muhammad Gondal, Manuel Wuthrich, Djordje Miladinovic, Francesco Locatello, Martin Breidt, Valentin Volchkov, Joel Akpo, Olivier Bachem, Bernhard Schölkopf and Stefan Bauer · 2019
Cited alongside, same era.
“An Introduction to Kolmogorov Complexity and Its Applications”
Ming Li and Paul Vitányi · 2019
Cited alongside, same era.
Later among the works it cites.
“Masked autoencoders are scalable vision learners”
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár and Ross Girshick · 2022
Later among the works it cites.
“Simplicial embeddings in self-supervised learning and downstream classification”, 2022
Samuel Lavoie, Christos Tsirigotis, Max Schwarzer, Ankit Vani, Michael Noukhovitch, Kenji Kawaguchi and Aaron Courville · 2022
Later among the works it cites.
“Invariant information bottleneck for domain generalization”
Bo Li, Yifei Shen, Yezhen Wang, Wenzhen Zhu, Dongsheng Li, Kurt Keutzer and Han Zhao · 2022
Later among the works it cites.
“The primacy bias in deep reinforcement learning”
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon and Aaron Courville · 2022
Later among the works it cites.
“Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation”, 2022
Aahlad Puli, Nitish Joshi, He He and Rajesh Ranganath · 2022
Later among the works it cites.
“Multi-label iterated learning for image classification with label ambiguity”
Sai Rajeswar, Pau Rodriguez, Soumye Singhal, David Vazquez and Aaron Courville · 2022
Later among the works it cites.
“Recipe for a General, Powerful, Scalable Graph Transformer”
Ladislav Rampášek, Mikhail Galkin, Vijay Dwivedi, Anh Luu, Guy Wolf and Dominique Beaini · 2022
Later among the works it cites.
“Better Supervisory Signals by Observing Learning Paths”
Yi Ren, Shangmin Guo and Danica. Sutherland · 2022
Later among the works it cites.
“Visual Representation Learning Does Not Generalize Strongly Within the Same Domain”
Lukas Schott, Julius Kügelgen, Frederik Träuble, Peter Gehler, Chris Russell, Matthias Bethge, Bernhard Schölkopf, Francesco Locatello and Wieland Brendel · 2022
Later among the works it cites.
“How robust are pre-trained models to distribution shift?”, 2022
Yuge Shi, Imant Daunhawer, Julia Vogt, Philip Torr and Amartya Sanyal · 2022
Later among the works it cites.
“Large-Scale Representation Learning on Graphs via Bootstrapping”
Shantanu Thakoor, Corentin Tallec, Mohammad Azar, Mehdi Azabou, Eva Dyer, Remi Munos, Petar Veličković and Michal Valko · 2022
Later among the works it cites.
“Compositional generalization in unsupervised compositional representation learning: A study on disentanglement and emergent language”
Zhenlin Xu, Marc Niethammer and Colin Raffel · 2022
Later among the works it cites.
“Fortuitous Forgetting in Connectionist Networks”
Hattie Zhou, Ankit Vani, Hugo Larochelle and Aaron Courville · 2022
Later among the works it cites.
“Image BERT Pre-training with Online Tokenizer”
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille and Tao Kong · 2022
Later among the works it cites.
“Probing Graph Representation”
Mohammad Akhondzadeh, Vijay Lingam and Aleksandar Bojchevski · 2023
Closest in time.
“Language modeling is compression”, 2023
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter and Joel Veness · 2023
Closest in time.
“Benchmarking graph neural networks”
Vijay Dwivedi, Chaitanya Joshi, Thomas Laurent, Yoshua Bengio and Xavier Bresson · 2023
Closest in time.
“Reinforced Self-Training (ReST) for Language Modeling”, 2023
Caglar Gulcehre, Tom Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat and Nando de Freitas · 2023
Closest in time.
“Self-refine: Iterative refinement with self-feedback”, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh and Peter Clark · 2023
Closest in time.
“Compression for AGI”, Stanford MLSys Seminar, 2023
Jack Rae · 2023
Closest in time.
“An observation on Generalization”, Simons Institute workshop on Large Language Models and Transformers, 2023
Ilya Sutskever · 2023
Closest in time.