Fetching the paper…
Reading the bibliography…
Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains.
MONet: Unsupervised Scene Decomposition and Representation
Christopher P. Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 1901
Earlier work this paper cites.
Multi-Object Representation Learning with Iterative Variational Inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 1903
Earlier work this paper cites.
Linjie Yang, Yuchen Fan, and Ning Xu · 1905
Earlier work this paper cites.
Entity Abstraction in Visual Model-based Reinforcement Learning
Rishi Veerapaneni, John D. Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine · 1910
Earlier work this paper cites.
Spatially invariant unsupervised object detection with convolutional neural networks
Eric Crawford and Joelle Pineau · 1911
Earlier work this paper cites.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei Efros · 2006
Earlier work this paper cites.
Vision meets Robotics: The KITTI Dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
Jordi Pont-Tuset, Pablo Arbeláez, Jonathan T. Barron, Ferran Marques, and Jitendra Malik · 2016
Earlier work this paper cites.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Discrete variational autoencoders
Jason Tyler Rolfe · 2017
Earlier work this paper cites.
Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Earlier work this paper cites.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Sjoerd van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Object-Centric Learning with Slot Attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Earlier work this paper cites.
SCALOR: Generative World Models with Scalable Object Representations
Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn · 2020
Earlier work this paper cites.
SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
Zhixuan Lin, Yi-Wu Fu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn · 2020
Earlier work this paper cites.
Self-supervised Visual Reinforcement Learning with Object-centric Representations
Andrii Zadaianchuk, Maximilian Seitzer, and Georg Martius · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
The 3rd large-scale video object segmentation challenge - video instance segmentation track, June 2021
Linjie Yang, Yuchen Fan, Yang Fu, and Ning Xu · 2021
Cited alongside, same era.
PARTS: Unsupervised segmentation with slots, attention and independence maximization
Daniel Zoran, Rishabh Kabra, Alexander Lerchner, and Danilo J. Rezende · 2021
Cited alongside, same era.
Benchmarking Unsupervised Object Representations for Video Sequences
Marissa A Weis, Kashyap Chitta, Yash Sharma, Wieland Brendel, Matthias Bethge, Andreas Geiger, and Alexander S. Ecker · 2021
Cited alongside, same era.
SIMONe: View-invariant, temporally-abstracted object representations via unsupervised video decomposition
Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess · 2021
Masked Siamese Networks for Label-Efficient Learning
Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael G. Rabbat, and Nicolas Ballas · 2022
Later among the works it cites.
Compositional Multi-object Reinforcement Learning with Linear Relation Networks
Davide Mambelli, Frederik Träuble, Stefan Bauer, Bernhard Schölkopf, and Francesco Locatello · 2022
Later among the works it cites.
Segmenting moving objects via an object-centric layered representation
Junyu Xie, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, and Francesco Locatello · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation
Laurynas Karazija, Iro Laina, and Christian Rupprecht · 2021
Cited alongside, same era.
GENESIS-V2: Inferring Unordered Object Representations without Iterative Refinement
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner · 2021
Cited alongside, same era.
An Empirical Study of Training Self-Supervised Vision Transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Sparsely changing latent states for prediction and planning in partially observable domains
Christian Gumbsch, Martin V Butz, and Georg Martius · 2021
Cited alongside, same era.
Conditional Object-centric Learning from Video
Thomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2022
Cited alongside, same era.
SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos
Gamaleldin Fathy Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael Curtis Mozer, and Thomas Kipf · 2022
Cited alongside, same era.
Sadra Safadoust and Fatma Güney · 2023
Closest in time.
Invariant slot attention: Object discovery with slot-centric reference frames
Ondrej Biza, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi, Gamaleldin F. Elsayed, Aravindh Mahendran, and Thomas Kipf · 2023
Closest in time.
Shepherding slots to objects: Towards stable and robust object-centric learning
Jinwoo Kim, Janghyuk Choi, Ho-Jin Choi, and Seon Joo Kim · 2023
Closest in time.
Differentiable mathematical programming for object-centric representation learning
Adeel Pervez, Phillip Lippe, and Efstratios Gavves · 2023
Closest in time.
Improving object-centric learning with query optimization
Baoxiong Jia, Yu Liu, and Siyuan Huang · 2023
Closest in time.
Jindong Jiang, Fei Deng, Gautam Singh, and Sungjin Ahn · 2023
Closest in time.
Object discovery from motion-guided tokens
Zhipeng Bao, Pavel Tokmakov, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert · 2023
Closest in time.
Semantics meets temporal correspondence: Self-supervised object-centric learning in videos
Rui Qian, Shuangrui Ding, Xian Liu, and Dahua Lin · 2023
Closest in time.
Self-supervised object-centric learning for videos
Görkay Aydemir, Weidi Xie, and Fatma Güney · 2023
Closest in time.
Unsupervised object learning via common fate
Matthias Tangemann, Steffen Schneider, Julius von Kügelgen, Francesco Locatello, Peter Gehler, Thomas Brox, Matthias Kümmerer, Matthias Bethge, and Bernhard Schölkopf · 2023
Closest in time.
Learning dynamic attribute-factored world models for efficient multi-object reinforcement learning
Fan Feng and Sara Magliacane · 2023
Closest in time.
Time does tell: Self-supervised time-tuning of dense image representations
Mohammadreza Salehi, Efstratios Gavves, Cees G.M. Snoek, and Yuki M. Asano · 2023
Closest in time.
DINOv2: Learning robust visual features without supervision, 2023
Maxime Oquab, Timothée Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Closest in time.