Fetching the paper…
Reading the bibliography…
Self-supervised methods have shown remarkable progress in learning high-level semantics and low-level temporal correspondence.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Towards understanding action recognition
Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J Black · 2013
Earlier work this paper cites.
Video segmentation by tracking many figure-ground segments
Fuxin Li, Taeyoung Kim, Ahmad Humayun, David Tsai, and James M Rehg · 2013
Earlier work this paper cites.
Segmentation of moving objects by long term video analysis
Peter Ochs, Jitendra Malik, and Thomas Brox · 2013
Earlier work this paper cites.
Learning a deep compact image representation for visual tracking
Naiyan Wang and Dit-Yan Yeung · 2013
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning to track at 100 fps with deep regression networks
David Held, Sebastian Thrun, and Silvio Savarese · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Scnet: Learning semantic correspondence
Kai Han, Rafael S Rezende, Bumsub Ham, Kwan-Yee K Wong, Minsu Cho, Cordelia Schmid, and Jean Ponce · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
End-to-end representation learning for correlation filter based tracking
Jack Valmadre, Luca Bertinetto, Joao Henriques, Andrea Vedaldi, and Philip HS Torr · 2017
Earlier work this paper cites.
Online adaptation of convolutional neural networks for video object segmentation
Paul Voigtlaender and Bastian Leibe · 2017
Earlier work this paper cites.
Transitive invariance for self-supervised visual representation learning
Xiaolong Wang, Kaiming He, and Abhinav Gupta · 2017
Earlier work this paper cites.
Flow-guided feature aggregation for video object detection
Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Earlier work this paper cites.
Self-supervised spatiotemporal feature learning via video rotation prediction
Longlong Jing, Xiaodong Yang, Jingen Liu, and Yingli Tian · 2018
Earlier work this paper cites.
Learning image representations by completing damaged jigsaw puzzles
Dahun Kim, Donghyeon Cho, Donggeun Yoo, and In So Kweon · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Actionflownet: Learning motion representation for action recognition
Joe Yue-Hei Ng, Jonghyun Choi, Jan Neumann, and Larry S Davis · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Earlier work this paper cites.
Tracking emerges by colorizing videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, and Kevin Murphy · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin · 2018
Earlier work this paper cites.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Cited alongside, same era.
Adaptive temporal encoding network for video instance-level human parsing
Qixian Zhou, Xiaodan Liang, Ke Gong, and Liang Lin · 2018
Cited alongside, same era.
Self-labelling via simultaneous clustering and representation learning
Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi · 2019
Cited alongside, same era.
Learning representations by maximizing mutual information across views
Philip Bachman, R Devon Hjelm, and William Buchwalter · 2019
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Self-supervision by prediction for object discovery in videos
Beril Besbinar and Pascal Frossard · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
Per-pixel classification is not all you need for semantic segmentation
Bowen Cheng, Alex Schwing, and Alexander Kirillov · 2021
Later among the works it cites.
Efficient iterative amortized inference for learning symmetric and disentangled multi-object representations
Patrick Emami, Pan He, Sanjay Ranka, and Anand Rangarajan · 2021
Later among the works it cites.
Genesis-v2: Inferring unordered object representations without iterative refinement
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The 2019 davis challenge on vos: Unsupervised multi-object segmentation
Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool · 2019
Cited alongside, same era.
Genesis: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R Kosiorek, Oiwi Parker Jones, and Ingmar Posner · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Cited alongside, same era.
Self-supervised learning for video correspondence flow
Zihang Lai and Weidi Xie · 2019
Cited alongside, same era.
Joint-task self-supervised learning for temporal correspondence
Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang · 2019
Cited alongside, same era.
Rvos: End-to-end recurrent network for video object segmentation
Carles Ventura, Miriam Bellver, Andreu Girbau, Amaia Salvador, Ferran Marques, and Xavier Giro-i Nieto · 2019
Cited alongside, same era.
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He · 2021
Later among the works it cites.
Simone: View-invariant, temporally-abstracted object representations via unsupervised video decomposition
Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey, Antonia Creswell, Matt Botvinick, Alexander Lerchner, and Chris Burgess · 2021
Later among the works it cites.
Conditional object-centric learning from video
Thomas Kipf, Gamaleldin F Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2021
Later among the works it cites.
Segmenting invisible moving objects
Hala Lamdouar, Weidi Xie, and Andrew Zisserman · 2021
Later among the works it cites.
Video instance segmentation with a propose-reduce paradigm
Huaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu, and Jiaya Jia · 2021
Later among the works it cites.
The emergence of objectness: Learning zero-shot segmentation from videos
Runtao Liu, Zhirong Wu, Stella Yu, and Stephen Lin · 2021
Later among the works it cites.
Enhancing self-supervised video representation learning via multi-level feature optimization
Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, and Weiyao Lin · 2021
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Later among the works it cites.
Dense contrastive learning for self-supervised visual pre-training
Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li · 2021
Later among the works it cites.
Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning
Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu · 2021
Later among the works it cites.
Rethinking self-supervised correspondence learning: A video frame-level similarity perspective
Jiarui Xu and Xiaolong Wang · 2021
Later among the works it cites.
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie · 2021
Later among the works it cites.
Learning pixel trajectories with multiscale contrastive random walks
Zhangxing Bian, Allan Jabri, Alexei A Efros, and Andrew Owens · 2022
Later among the works it cites.
Guess what moves: unsupervised video and image segmentation by anticipating motion
Subhabrata Choudhury, Laurynas Karazija, Iro Laina, Andrea Vedaldi, and Christian Rupprecht · 2022
Later among the works it cites.
Motion-aware contrastive video representation learning via foreground-background merging
Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, and Hongkai Xiong · 2022
Later among the works it cites.
Dual contrastive learning for spatio-temporal representation
Shuangrui Ding, Rui Qian, and Hongkai Xiong · 2022
Later among the works it cites.
Motion-inductive self-supervised object discovery in videos
Shuangrui Ding, Weidi Xie, Yabo Chen, Rui Qian, Xiaopeng Zhang, Hongkai Xiong, and Qi Tian · 2022
Later among the works it cites.
Savi++: Towards end-to-end object-centric learning from real-world videos
Gamaleldin F Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael C Mozer, and Thomas Kipf · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Semantic-aware fine-grained correspondence
Yingdong Hu, Renhao Wang, Kaifeng Zhang, and Yang Gao · 2022
Later among the works it cites.
Unsupervised object-centric learning with bi-level optimized query slot attention
Baoxiong Jia, Yu Liu, and Siyuan Huang · 2022
Later among the works it cites.
Static and dynamic concepts for self-supervised video representation learning
Rui Qian, Shuangrui Ding, Xian Liu, and Dahua Lin · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, et al · 2022
Later among the works it cites.
Segmenting moving objects via an object-centric layered representation
Junyu Xie, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Prune spatio-temporal tokens by semantic-aware temporal accumulation
Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang, Rui Qian, Hongkai Xiong, and Qi Tian · 2023
Closest in time.