Fetching the paper…
Reading the bibliography…
The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions.
Robust estimation of a location parameter
Peter J. Huber · 1964
Earlier work this paper cites.
Objective criteria for the evaluation of clustering methods
William M Rand · 1971
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie · 1985
Earlier work this paper cites.
Serial and parallel processing of visual feature conjunctions
K. Nakayama and G. H. Silverman · 1986
Earlier work this paper cites.
Principles of object perception
E. S. Spelke · 1990
Earlier work this paper cites.
Motion coherence and conjunction search: Implications for guided search theory
J. Driver, P. McLeod, and Z. Dienes · 1992
Earlier work this paper cites.
The reviewing of object files: Object-specific integration of information
Daniel Kahneman, Anne Treisman, and Brian J Gibbs · 1992
Earlier work this paper cites.
Visual search for type of motion is based on simple motion primitives
Todd S Horowitz, Jeremy M Wolfe, Jennifer S DiMase, and Sarah B Klieger · 2007
Earlier work this paper cites.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler · 2007
Earlier work this paper cites.
Object segmentation by long term analysis of point trajectories
Thomas Brox and Jitendra Malik · 2010
Earlier work this paper cites.
Object segmentation in video: a hierarchical variational approach for turning point trajectories into dense regions
Peter Ochs and Thomas Brox · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Earlier work this paper cites.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
The 2017 DAVIS challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Cited alongside, same era.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Cited alongside, same era.
Deep image prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2018
Cited alongside, same era.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Sjoerd van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
MONet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Entity abstraction in visual model-based reinforcement learning
Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine · 2020
Later among the works it cites.
SDC-depth: Semantic divide-and-conquer network for monocular depth estimation
Lijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin, and Huchuan Lu · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Track, check, repeat: An EM approach to unsupervised tracking
Adam W Harley, Yiming Zuo, Jing Wen, Ayush Mangal, Shubhankar Potdar, Ritwick Chaudhry, and Katerina Fragkiadaki · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The 2019 DAVIS challenge on vos: Unsupervised multi-object segmentation
Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool · 2019
Cited alongside, same era.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Learning physical graph representations from visual scenes
Daniel M Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei, Jiajun Wu, Joshua B Tenenbaum, et al · 2020
Cited alongside, same era.
Tao: A large-scale benchmark for tracking any object
Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan · 2020
Cited alongside, same era.
Google scanned objects, 2020
Google Research · 2020
Cited alongside, same era.
SIMONe: View-invariant, temporally-abstracted object representations via unsupervised video decomposition
Rishabh Kabra, Daniel Zoran, Loic Matthey Goker Erdogan, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess · 2021
Later among the works it cites.
MDETR – Modulated Detection for End-to-End Multi-Modal Understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Later among the works it cites.
ClevrTex: A texture-rich benchmark for unsupervised multi-object segmentation
Laurynas Karazija, Iro Laina, and Christian Rupprecht · 2021
Later among the works it cites.
TrackFormer: Multi-Object Tracking with Transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Deep learning for monocular depth estimation: A review
Yue Ming, Xuyang Meng, Chunxiao Fan, and Hui Yu · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
Toward causal representation learning
Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio · 2021
Later among the works it cites.
Unsupervised object detection with lidar clues
Hao Tian, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang, and Xizhou Zhu · 2021
Later among the works it cites.
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie · 2021
Later among the works it cites.
Parts: Unsupervised segmentation with slots, attention and independence maximization
Daniel Zoran, Rishabh Kabra, Alexander Lerchner, and Danilo J Rezende · 2021
Later among the works it cites.
Discovering objects that can move
Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert · 2022
Closest in time.
Object discovery and representation networks
Olivier J Hénaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelović · 2022
Closest in time.
Conditional object-centric learning from video
Thomas Kipf, Gamaleldin F Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2022
Closest in time.
Drive&segment: Unsupervised semantic segmentation of urban scenes via cross-modal distillation
Antonin Vobecky, David Hurych, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, and Josef Sivic · 2022
Closest in time.
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Closest in time.
Slot-vps: Object-centric representation learning for video panoptic segmentation
Yi Zhou, Hui Zhang, Hana Lee, Shuyang Sun, Pingjun Li, Yangguang Zhu, ByungIn Yoo, Xiaojuan Qi, and Jae-Joon Han · 2022
Closest in time.