Fetching the paper…
Reading the bibliography…
We introduce Cosmos, a framework for object-centric world modeling that is designed for compositional generalization (CompGen), i.e., high performance on unseen input scenes obtained through the composition of known visual "atoms." The central insight behind Cosmos is the use of a novel form of neurosymbolic grounding.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes
Nicholas Watters, Loic Matthey, Christopher P Burgess, and Alexander Lerchner · 1901
Earlier work this paper cites.
Nicholas Watters, Loic Matthey, Matko Bosnjak, Christopher P Burgess, and Alexander Lerchner · 1905
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnad · 1990
Earlier work this paper cites.
Plannable approximations to mdp homomorphisms: Equivariance under actions
Elise Van der Pol, Thomas Kipf, Frans A Oliehoek, and Max Welling · 2002
Earlier work this paper cites.
An algebraic approach to abstraction in reinforcement learning
Balaraman Ravindran · 2004
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Houdini: Lifelong learning as program synthesis
Lazar Valkov, Dipak Chaudhari, Akash Srivastava, Charles Sutton, and Swarat Chaudhuri · 2018
Earlier work this paper cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Earlier work this paper cites.
Contrastive learning of structured world models
Thomas Kipf, Elise Van der Pol, and Max Welling · 2019
Earlier work this paper cites.
Program-guided image manipulators
Jiayuan Mao, Xiuming Zhang, Yikai Li, William T Freeman, Joshua B Tenenbaum, and Jiajun Wu · 2019
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet · 2020
Cited alongside, same era.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Cited alongside, same era.
Learning differentiable programs with admissible neural heuristics
Ameesh Shah, Eric Zhan, Jennifer J Sun, Abhinav Verma, Yisong Yue, and Swarat Chaudhuri · 2020
Cited alongside, same era.
Entity abstraction in visual model-based reinforcement learning
Rishi Veerapaneni, John D. Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine · 2020
Cited alongside, same era.
PDSketch: Integrated Domain Programming, Learning, and Planning
Jiayuan Mao, Tomas Lozano-Perez, Joshua B. Tenenbaum, and Leslie Pack Kaelbing · 2022
Later among the works it cites.
Learning to compose soft prompts for compositional zero-shot learning
Nihal V Nayak, Peilin Yu, and Stephen H Bach · 2022
Later among the works it cites.
Learning symmetric embeddings for equivariant world models
Jung Yeon Park, Ondrej Biza, Linfeng Zhao, Jan Willem van de Meent, and Robin Walters · 2022
Later among the works it cites.
Toward compositional generalization in object-oriented world modeling
Linfeng Zhao, Lingzhi Kong, Robin Walters, and Lawson LS Wong · 2022
Later among the works it cites.
Hierarchical abstraction for combinatorial generalization in object rearrangement
Michael Chang, Alyssa L. Dayan, Franziska Meier, Thomas L. Griffiths, Sergey Levine, and Amy Zhang · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robin Walters, Jinxi Li, and Rose Yu · 2020
Cited alongside, same era.
Neural production systems
Anirudh Goyal, Aniket Didolkar, Nan Rosemary Ke, Charles Blundell, Philippe Beaudoin, Nicolas Heess, Michael Mozer, and Yoshua Bengio · 2021
Cited alongside, same era.
Task programming: Learning data efficient behavior representations
Jennifer J Sun, Ann Kennedy, Eric Zhan, David J Anderson, Yisong Yue, and Pietro Perona · 2021
Cited alongside, same era.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Nan Rosemary Ke, Aniket Didolkar, Sarthak Mittal, Anirudh Goyal, Guillaume Lajoie, Stefan Bauer, Danilo Rezende, Yoshua Bengio, Michael Mozer, and Christopher Pal · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Unsupervised learning of neurosymbolic encoders
Eric Zhan, Jennifer J. Sun, Ann Kennedy, Yisong Yue, and Swarat Chaudhuri · 2021
Cited alongside, same era.
Object representations as fixed points: Training iterative refinement algorithms with implicit differentiation
Michael Chang, Tom Griffiths, and Sergey Levine · 2022
Cited alongside, same era.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
Ns3d: Neuro-symbolic grounding of 3d objects and relations
Joy Hsu, Jiayuan Mao, and Jiajun Wu · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Closest in time.
From perception to programs: regularize, overparameterize, and amortize
Hao Tang and Kevin Ellis · 2023
Closest in time.
Slotformer: Unsupervised visual dynamics simulation with object-centric models, 2023
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2023
Closest in time.