Fetching the paper…
Reading the bibliography…
Object discovery -- separating objects from the background without manual labels -- is a fundamental open challenge in computer vision.
Vector quantization
Robert Gray · 1984
Earlier work this paper cites.
The reviewing of object files: Object-specific integration of information
Daniel Kahneman, Anne Treisman, and Brian J Gibbs · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient graph-based image segmentation
Pedro F Felzenszwalb and Daniel P Huttenlocher · 2004
Earlier work this paper cites.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler · 2007
Earlier work this paper cites.
Introduction to information retrieval
Hinrich Schütze, Christopher D Manning, and Prabhakar Raghavan · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Principles of Gestalt psychology
Kurt Koffka · 2013
Earlier work this paper cites.
Multiscale combinatorial grouping
Pablo Arbeláez, Jordi Pont-Tuset, Jonathan T Barron, Ferran Marques, and Jitendra Malik · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Delving deeper into convolutional networks for learning video representations
Nicolas Ballas, Li Yao, Chris Pal, and Aaron Courville · 2016
Earlier work this paper cites.
Tagger: Deep unsupervised perceptual grouping
Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Discrete variational autoencoders
Jason Tyler Rolfe · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
MONet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Earlier work this paper cites.
Towards segmenting anything that moves
Achal Dave, Pavel Tokmakov, and Deva Ramanan · 2019
Earlier work this paper cites.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Learning two-view correspondences and geometry using order-aware network
Jiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao, Lei Zhou, Tianwei Shen, Yurong Chen, Long Quan, and Hongen Liao · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Decomposing 3D scenes into objects via unsupervised volume segmentation
Karl Stelzner, Kristian Kersting, and Adam R Kosiorek · 2021
Later among the works it cites.
SMURF: Self-teaching multi-frame unsupervised RAFT with full-image warping
Austin Stone, Daniel Maurer, Alper Ayvaci, Anelia Angelova, and Rico Jonschkowski · 2021
Later among the works it cites.
Unsupervised semantic segmentation by contrasting object mask proposals
Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Later among the works it cites.
Unsupervised foreground extraction via deep region competition
Peiyu Yu, Sirui Xie, Xiaojian Ma, Yixin Zhu, Ying Nian Wu, and Song-Chun Zhu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GENESIS: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R Kosiorek, Oiwi Parker Jones, and Ingmar Posner · 2020
Cited alongside, same era.
3d packing for self-supervised monocular depth estimation
Rares Guizilini, Vitor abd Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon · 2020
Cited alongside, same era.
Unsupervised object-centric video generation and decomposition in 3D
Paul Henderson and Christoph H Lampert · 2020
Cited alongside, same era.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei Efros · 2020
Cited alongside, same era.
SCALOR: Generative world models with scalable object representations
Jindong Jiang, Sepehr Janghorbani, Gerard De Melo, and Sungjin Ahn · 2020
Cited alongside, same era.
Space: Unsupervised object-oriented scene representation via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn · 2020
Cited alongside, same era.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Later among the works it cites.
Discovering objects that can move
Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert · 2022
Later among the works it cites.
Learning pixel trajectories with multiscale contrastive random walks
Zhangxing Bian, Allan Jabri, Alexei A Efros, and Andrew Owens · 2022
Later among the works it cites.
Savi++: Towards end-to-end object-centric learning from real-world videos
Gamaleldin F Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael C Mozer, and Thomas Kipf · 2022
Later among the works it cites.
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2022
Later among the works it cites.
Unsupervised multi-object segmentation by predicting probable motion patterns
Laurynas Karazija, Subhabrataand Choudhury, Iro Laina, Christian Rupprecht, and Andrea Vedaldi · 2022
Later among the works it cites.
Conditional object-centric learning from video
Thomas Kipf, Gamaleldin F Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2022
Later among the works it cites.
Trackformer: Multi-object tracking with transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Object scene representation transformer
Mehdi SM Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste, Filip Pavetić, Mario Lučić, Leonidas J Guibas, Klaus Greff, and Thomas Kipf · 2022
Later among the works it cites.
Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations
Mehdi SM Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani Vora, Mario Lučić, Daniel Duckworth, Alexey Dosovitskiy, et al · 2022
Later among the works it cites.
Simple unsupervised object-centric learning for complex and naturalistic videos
Gautam Singh, Yi-Fu Wu, and Sungjin Ahn · 2022
Later among the works it cites.
Cross-view transformers for real-time map-view semantic segmentation
Brady Zhou and Philipp Krähenbühl · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, et al · 2023
Closest in time.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2023
Closest in time.