Fetching the paper…
Reading the bibliography…
Although DALL-E has shown an impressive ability of composition-based systematic generalization in image generation, it requires the dataset of text-image pairs and the compositionality is provided by the text.
Exploiting spatial invariance for scalable unsupervised object tracking
Eric Crawford and Joelle Pineau · 1911
Earlier work this paper cites.
Markov chain sampling methods for dirichlet process mixture models
Radford M Neal · 2000
Earlier work this paper cites.
Vision as bayesian inference: analysis by synthesis?
Alan Yuille and Daniel Kersten · 2006
Earlier work this paper cites.
On the binding problem in artificial neural networks
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2012
Earlier work this paper cites.
On the binding problem in artificial neural networks
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2012
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi · 2014
Earlier work this paper cites.
Deep convolutional inverse graphics network
Tejas D Kulkarni, Will Whitney, Pushmeet Kohli, and Joshua B Tenenbaum · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Xinchen Yan, and Honglak Lee · 2015
Earlier work this paper cites.
Infogan: interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Adversarial feature learning
Jeff Donahue, Philipp Krähenbühl, and Trevor Darrell · 2016
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Generating images part by part with composite generative adversarial networks
Hanock Kwak and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
An algorithm for online k-means clustering
Edo Liberty, Ram Sriharsha, and Maxim Sviridenko · 2016
Earlier work this paper cites.
Attribute2image: Conditional image generation from visual attributes
Xinchen Yan, Jimei Yang, Kihyuk Sohn, and Honglak Lee · 2016
Earlier work this paper cites.
The cognitive map in humans: spatial navigation and beyond
Russell A Epstein, Eva Zita Patai, Joshua B Julian, and Hugo J Spiers · 2017
Earlier work this paper cites.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Earlier work this paper cites.
Denoising criterion for variational auto-encoding framework
Daniel Im Jiwoong Im, Sungjin Ahn, Roland Memisevic, and Yoshua Bengio · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Variational inference of disentangled latent concepts from unlabeled observations
Abhishek Kumar, Prasanna Sattigeri, and Avinash Balakrishnan · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Systematic generalization: what is required and can it be learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville · 2018
Earlier work this paper cites.
What is a cognitive map? organizing knowledge for flexible behavior
Timothy EJ Behrens, Timothy H Muller, James CR Whittington, Shirley Mark, Alon B Baram, Kimberly L Stachenfeld, and Zeb Kurth-Nelson · 2018
Earlier work this paper cites.
3d shapes dataset
Chris Burgess and Hyunjik Kim · 2018
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Tian Qi Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud · 2018
Earlier work this paper cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Oliver Groth, Fabian B. Fuchs, Ingmar Posner, and Andrea Vedaldi · 2018
Cited alongside, same era.
Scan: Learning hierarchical compositional visual concepts
Irina Higgins, Nicolas Sonnerat, Loic Matthey, Arka Pal, Christopher P Burgess, Matko Bosnjak, Murray Shanahan, Matthew Botvinick, Demis Hassabis, and Alexander Lerchner · 2018
Cited alongside, same era.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Cited alongside, same era.
Hyunjik Kim and Andriy Mnih · 2018
Cited alongside, same era.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Learning canonical representations for scene graph to image generation
Roei Herzig, Amir Bar, Huijuan Xu, Gal Chechik, Trevor Darrell, and Amir Globerson · 2020
Later among the works it cites.
Generative neurosymbolic machines
Jindong Jiang and Sungjin Ahn · 2020
Later among the works it cites.
Towards unsupervised learning of generative models for 3d controllable image synthesis
Yiyi Liao, Katja Schwarz, Lars Mescheder, and Andreas Geiger · 2020
Later among the works it cites.
Object-centric learning with slot attention, 2020
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Blockgan: Learning 3d object-aware scene representations from unlabelled images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Specifying object attributes and relations in interactive scene generation
Oron Ashual and Lior Wolf · 2019
Cited alongside, same era.
Emergence of object segmentation in perturbed generative models
Adam Bielski and Paolo Favaro · 2019
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Unsupervised object segmentation by redrawing
Mickaël Chen, Thierry Artières, and Ludovic Denoyer · 2019
Cited alongside, same era.
Large scale adversarial representation learning
Jeff Donahue and Karen Simonyan · 2019
Cited alongside, same era.
Thu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang, and Niloy Mitra · 2020
Later among the works it cites.
Investigating object compositionality in generative adversarial networks
Sjoerd van Steenkiste, Karol Kurach, Jürgen Schmidhuber, and Sylvain Gelly · 2020
Later among the works it cites.
Towards causal generative scene models via competition of experts
Julius von Kügelgen, Ivan Ustyuzhaninov, Peter V. Gehler, Matthias Bethge, and Bernhard Schölkopf · 2020
Later among the works it cites.
Learning to manipulate individual objects in an image
Yanchao Yang, Yutong Chen, and Stefano Soatto · 2020
Later among the works it cites.
Compositional video synthesis with action graphs
Amir Bar, Roi Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, and Amir Globerson · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Closest in time.
Using latent space regression to analyze and leverage compositionality in gans
Lucy Chai, Jonas Wulff, and Phillip Isola · 2021
Closest in time.
Generative scene graph networks
Fei Deng, Zhuo Zhi, Donghun Lee, and Sungjin Ahn · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Genesis-v2: Inferring unordered object representations without iterative refinement
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner · 2021
Closest in time.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2021
Closest in time.
Bitmoji dataset
Romain Graux · 2021
Closest in time.
How to represent part-whole hierarchies in a neural network
Geoffrey E. Hinton · 2021
Closest in time.
Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess · 2021
Closest in time.
Clevrtex: A texture-rich benchmark for unsupervised multi-object segmentation
Laurynas Karazija, Iro Laina, and Christian Rupprecht · 2021
Closest in time.
Transformers with competitive ensembles of independent mechanisms
Alex Lamb, Di He, Anirudh Goyal, Guolin Ke, Chien-Feng Liao, Mirco Ravanelli, and Yoshua Bengio · 2021
Closest in time.
Giraffe: Representing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger · 2021
Closest in time.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
Information-theoretic segmentation by inpainting error maximization
Pedro Savarese, Sunnie S. Y. Kim, Michael Maire, Greg Shakhnarovich, and David McAllester · 2021
Closest in time.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Douze Matthijs, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Closest in time.
Big gans are watching you: Towards unsupervised object segmentation with off-the-shelf generative models
Andrey Voynov, Stanislav Morozov, and Artem Babenko · 2021
Closest in time.
Generative video transformer: Can objects be the words?
Yi-Fu Wu, Jaesik Yoon, and Sungjin Ahn · 2021
Closest in time.
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie · 2021
Closest in time.