Fetching the paper…
Reading the bibliography…
Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built.
The reviewing of object files: Object-specific integration of information
Daniel Kahneman, Anne Treisman, and Brian J Gibbs · 1992
Earlier work this paper cites.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Energy-based geometric multi-model fitting
Hossam Isack and Yuri Boykov · 2012
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Video segmentation by non-local consensus voting
Alon Faktor and Michal Irani · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Dense optical flow prediction from a static image
Jacob Walker, Abhinav Gupta, and Martial Hebert · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Geoffrey E Hinton, et al · 2016
Earlier work this paper cites.
Tagger: Deep unsupervised perceptual grouping
Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Structural-RNN: Deep learning on spatio-temporal graphs
Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2016
Earlier work this paper cites.
One-shot video object segmentation
Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset, Laura Leal-Taixé, Daniel Cremers, and Luc Van Gool · 2017
Earlier work this paper cites.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
The 2017 DAVIS challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Matrix capsules with EM routing
Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst · 2018
Earlier work this paper cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Earlier work this paper cites.
Premvos: Proposal-generation, refinement and merging for video object segmentation
Jonathon Luiten, Paul Voigtlaender, and Bastian Leibe · 2018
Earlier work this paper cites.
Relational recurrent neural networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, and Timothy Lillicrap · 2018
Earlier work this paper cites.
PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Cited alongside, same era.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Sjoerd van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
MONet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
The 2019 DAVIS challenge on vos: Unsupervised multi-object segmentation
Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool · 2019
Cited alongside, same era.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei A Efros · 2020
Later among the works it cites.
SCALOR: Generative world models with scalable object representations
Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn · 2020
Later among the works it cites.
Contrastive learning of structured world models
Thomas Kipf, Elise van der Pol, and Max Welling · 2020
Later among the works it cites.
SPACE: Unsupervised object-oriented scene representation via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn · 2020
Later among the works it cites.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spatially invariant unsupervised object detection with convolutional neural networks
Eric Crawford and Joelle Pineau · 2019
Cited alongside, same era.
End-to-end recurrent multi-object tracking and trajectory prediction with relational reasoning
Fabian B Fuchs, Adam R Kosiorek, Li Sun, Oiwi Parker Jones, and Ingmar Posner · 2019
Cited alongside, same era.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 2019
Cited alongside, same era.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Spatio-temporal action graph networks
Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson, and Trevor Darrell · 2019
Cited alongside, same era.
Joint-task self-supervised learning for temporal correspondence
Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang · 2019
Cited alongside, same era.
Sara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E Hinton, and David J Fleet · 2020
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine · 2020
Later among the works it cites.
Unmasking the inductive biases of unsupervised object representations for video sequences
Marissa A Weis, Kashyap Chitta, Yash Sharma, Wieland Brendel, Matthias Bethge, Andreas Geiger, and Alexander S Ecker · 2020
Later among the works it cites.
CLEVRER: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2020
Later among the works it cites.
A transductive approach for video object segmentation
Yizhuo Zhang, Zhirong Wu, Houwen Peng, and Stephen Lin · 2020
Later among the works it cites.
Compositional video synthesis with action graphs
Amir Bar, Roei Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, and Amir Globerson · 2021
Closest in time.
Blender - a 3D modelling and rendering package
Blender Online Community · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Closest in time.
Grounding physical concepts of objects and events through dynamic visual reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B Tenenbaum, and Chuang Gan · 2021
Closest in time.
Unsupervised object-based transition models for 3D partially observable environments
Antonia Creswell, Rishabh Kabra, Chris Burgess, and Murray Shanahan · 2021
Closest in time.
Kubric: A data generation pipeline for creating semi-realistic synthetic multi-object videos, 2021
Klaus Greff, Andrea Tagliasacchi, Derek Liu, Issam Laradji, Or Litany, and Luca Prasso · 2021
Closest in time.
Track, check, repeat: An EM approach to unsupervised tracking
Adam W Harley, Yiming Zuo, Jing Wen, Ayush Mangal, Shubhankar Potdar, Ritwick Chaudhry, and Katerina Fragkiadaki · 2021
Closest in time.
Rishabh Kabra, Daniel Zoran, Loic Matthey Goker Erdogan, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess · 2021
Closest in time.
MDETR – Modulated Detection for End-to-End Multi-Modal Understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Closest in time.
Clevrtex: A texture-rich benchmark for unsupervised multi-object segmentation, 2021
Laurynas Karazija, Iro Laina, and Christian Rupprecht · 2021
Closest in time.
TrackFormer: Multi-Object Tracking with Transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer · 2021
Closest in time.
Decomposing 3D scenes into objects via unsupervised volume segmentation
Karl Stelzner, Kristian Kersting, and Adam R Kosiorek · 2021
Closest in time.
SMURF: Self-teaching multi-frame unsupervised RAFT with full-image warping
Austin Stone, Daniel Maurer, Alper Ayvaci, Anelia Angelova, and Rico Jonschkowski · 2021
Closest in time.
Towards general purpose vision systems
Aniruddha Kembhavi Derek Hoiem Tanmay Gupta, Amita Kamath · 2021
Closest in time.
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie · 2021
Closest in time.
TubeR: Tube-Transformer for Action Detection
Jiaojiao Zhao, Xinyu Li, Chunhui Liu, Shuai Bing, Hao Chen, Cees GM Snoek, and Joseph Tighe · 2021
Closest in time.