Fetching the paper…
Reading the bibliography…
Spatially dense self-supervised learning is a rapidly growing problem domain with promising applications for unsupervised segmentation and pretraining for dense downstream tasks.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Large displacement optical flow: Descriptor matching in variational motion estimation
Thomas Brox and Jitendra Malik · 2011
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Joint optical flow and temporally consistent semantic segmentation
Junhwa Hur and Stefan Roth · 2016
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Semantic video cnns through representation warping
Raghudeep Gadde, Varun Jampani, and Peter V Gehler · 2017
Earlier work this paper cites.
Surveillance video parsing with single frame supervision
Si Liu, Changhu Wang, Ruihe Qian, Han Yu, Renda Bao, and Yao Sun · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Deep feature flow for video recognition
Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Earlier work this paper cites.
Low-latency video semantic segmentation
Yule Li, Jianping Shi, and Dahua Lin · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Earlier work this paper cites.
Youtube-vos: A large-scale video object segmentation benchmark
Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas Huang · 2018
Earlier work this paper cites.
Dynamic video segmentation network
Yu-Syuan Xu, Tsu-Jui Fu, Hsuan-Kung Yang, and Chun-Yi Lee · 2018
Earlier work this paper cites.
Accel: A corrective fusion network for efficient semantic segmentation on video
Samvit Jain, Xin Wang, and Joseph E Gonzalez · 2019
Earlier work this paper cites.
Invariant information clustering for unsupervised image classification and segmentation
Xu Ji, Joao F Henriques, and Andrea Vedaldi · 2019
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Earlier work this paper cites.
Joint-task self-supervised learning for temporal correspondence
Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A Efros · 2019
Earlier work this paper cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Earlier work this paper cites.
Self-labelling via simultaneous clustering and representation learning
Yuki M. Asano, Christian Rupprecht, and Andrea Vedaldi · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Watching the world go by: Representation learning from unlabeled videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, and Ali Farhadi · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Temporally distributed networks for fast video semantic segmentation
Ping Hu, Fabian Caba, Oliver Wang, Zhe Lin, Stan Sclaroff, and Federico Perazzi · 2020
Epic-kitchens visor benchmark: Video segmentations and object relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen · 2022
Later among the works it cites.
Unsupervised semantic segmentation by distilling feature correspondences
Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T Freeman · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Object discovery and representation networks
Olivier J Hénaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelović · 2022
Later among the works it cites.
Video k-net: A simple, strong, and unified baseline for video segmentation
Xiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen, Guangliang Cheng, Yunhai Tong, and Chen Change Loy · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei Efros · 2020
Cited alongside, same era.
Efficient semantic video segmentation with per-frame inference
Yifan Liu, Chunhua Shen, Changqian Yu, and Jingdong Wang · 2020
Cited alongside, same era.
Learning video object segmentation from unlabeled videos
Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David J Crandall, and Steven CH Hoi · 2020
Cited alongside, same era.
Scan: Learning to classify images without labels
Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Marc Proesmans, and Luc Van Gool · 2020
Cited alongside, same era.
Dense unsupervised learning for video segmentation
Nikita Araslanov, Simone Schaub-Meyer, and Stefan Roth · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
An empirical study of training self-supervised vision transformers
Xinlei Chen*, Saining Xie*, and Kaiming He · 2021
Cited alongside, same era.
Self-supervised video pretraining yields strong image representations
Nikhil Parthasarathy, SM Eslami, João Carreira, and Olivier J Hénaff · 2022
Later among the works it cites.
Self-supervised video transformer
Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, and Michael S Ryoo · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, et al · 2022
Later among the works it cites.
Coarse-to-fine feature mining for video semantic segmentation
Guolei Sun, Yun Liu, Henghui Ding, Thomas Probst, and Luc Van Gool · 2022
Later among the works it cites.
How severe is benchmark-sensitivity in video self-supervised learning?
Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, and Cees GM Snoek · 2022
Later among the works it cites.
Look before you match: Instance understanding matters in video object segmentation
Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Chuanxin Tang, Xiyang Dai, Yucheng Zhao, Yujia Xie, Lu Yuan, and Yu-Gang Jiang · 2022
Later among the works it cites.
Patch-level representation learning for self-supervised vision transformers
Sukmin Yun, Hankook Lee, Jaehyung Kim, and Jinwoo Shin · 2022
Later among the works it cites.
Unsupervised semantic segmentation with self-supervised object-centric representations
Andrii Zadaianchuk, Matthaeus Kleindessner, Yi Zhu, Francesco Locatello, and Thomas Brox · 2022
Later among the works it cites.
ibot: Image bert pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong · 2022
Later among the works it cites.
Self-supervised learning of object parts for semantic segmentation
Adrian Ziegler and Yuki M Asano · 2022
Later among the works it cites.
Time to augment self-supervised visual representation learning
Arthur Aubret, Markus R. Ernst, Céline Teulière, and Jochen Triesch · 2023
Closest in time.
Towards in-context scene understanding
Ivana Balažević, David Steiner, Nikhil Parthasarathy, Relja Arandjelović, and Olivier J Hénaff · 2023
Closest in time.
Visual representation learning from unlabeled video using contrastive masked autoencoders
Jefferson Hernandez, Ruben Villegas, and Vicente Ordonez · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Closest in time.
Croc: Cross-view online clustering for dense visual representation learning
Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars, and Jean-Philippe Thiran · 2023
Closest in time.
Cut and learn for unsupervised object detection and instance segmentation
Xudong Wang, Rohit Girdhar, Stella X Yu, and Ishan Misra · 2023
Closest in time.
Patch-level contrasting without patch correspondence for accurate and dense contrastive representation learning
Shaofeng Zhang, Feng Zhu, Rui Zhao, and Junchi Yan · 2023
Closest in time.
Optical flow boosts unsupervised localization and segmentation
Xinyu Zhang and Abdeslam Boularias · 2023
Closest in time.
Boosting video object segmentation via space-time correspondence learning
Yurong Zhang, Liulei Li, Wenguan Wang, Rong Xie, Li Song, and Wenjun Zhang · 2023
Closest in time.