Fetching the paper…
Reading the bibliography…
Despite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow.
High accuracy optical flow estimation based on a theory for warping
Thomas Brox, Andrés Bruhn, Nils Papenberg, and Joachim Weickert · 2004
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
High-resolution stereo datasets with subpixel-accurate ground truth
Daniel Scharstein, Heiko Hirschmüller, York Kitajima, Greg Krathwohl, Nera Nešić, Xi Wang, and Porter Westling · 2014
Earlier work this paper cites.
A quantitative analysis of current practices in optical flow estimation and the principles behind them
Deqing Sun, Stefan Roth, and Michael J Black · 2014
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox · 2015
Earlier work this paper cites.
Object scene flow for autonomous vehicles
Moritz Menze and Andreas Geiger · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Deepmatching: Hierarchical deformable dense matching
Jerome Revaud, Philippe Weinzaepfel, Zaid Harchaoui, and Cordelia Schmid · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness
Jason J Yu, Adam W Harley, and Konstantinos G Derpanis · 2016
Earlier work this paper cites.
Generic 3D representation via pose estimation and matching
Amir R Zamir, Tilman Wekel, Pulkit Agrawal, Colin Wei, Jitendra Malik, and Silvio Savarese · 2016
Earlier work this paper cites.
Colorful Image Colorization
Richard Zhang, Phillip Isola, and Alexei A. Efros · 2016
Earlier work this paper cites.
Unsupervised monocular depth estimation with left-right consistency
Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow · 2017
Earlier work this paper cites.
End-to-end learning of geometry and context for deep stereo regression
A. Kendall, H. Martirosyan, S. Dasgupta, P. Henry, R. Kennedy, A. Bachrach, and A. Bry · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Scene parsing through ADE20K Dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Unsupervised learning of stereo matching
Chao Zhou, Hong Zhang, Xiaoyong Shen, and Jiaya Jia · 2017
Earlier work this paper cites.
Pyramid stereo matching network
J. Chang and Y. Chen · 2018
Earlier work this paper cites.
Unsupervised Representation Learning by Predicting Image Rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Earlier work this paper cites.
Left-right comparative recurrent model for stereo matching
Zequn Jie, Pengfei Wang, Yonggen Ling, Bo Zhao, Yunchao Wei, Jiashi Feng, and Wei Liu · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction
Sameh Khamis, Sean Fanello, Christoph Rhemann, Adarsh Kowdle, Julien Valentin, and Shahram Izadi · 2018
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Self-attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Attention Augmented Convolutional Networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Digging into self-supervised monocular depth estimation
Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow · 2019
Earlier work this paper cites.
Selflow: Self-supervised learning of optical flow
Pengpeng Liu, Michael Lyu, Irwin King, and Jia Xu · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Multi-level context ultra-aggregation for stereo matching
Guang-Yu Nie, Ming-Ming Cheng, Yun Liu, Zhengfa Liang, Deng-Ping Fan, Yue Liu, and Yongtian Wang · 2019
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Models matter, so does training: An empirical study of cnns for optical flow estimation
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2019
Cited alongside, same era.
Hierarchical discrete distribution decomposition for match density estimation
Zhichao Yin, Trevor Darrell, and Fisher Yu · 2019
Cited alongside, same era.
Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Cited alongside, same era.
Dense Contrastive Learning for Self-Supervised Visual Pre-Training
Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li · 2021
Later among the works it cites.
Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning
Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu · 2021
Later among the works it cites.
Optical flow and scene flow estimation: A survey
Mingliang Zhai, Xuezhi Xiang, Ning Lv, and Xiangdong Kong · 2021
Later among the works it cites.
Separable flow: Learning motion cost volumes for optical flow estimation
Feihu Zhang, Oliver J. Woodford, Victor Prisacariu, and Philip H. S. Torr · 2021
Later among the works it cites.
Masked siamese networks for label-efficient learning
Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas · 2022
Closest in time.
MultiMAE: Multi-modal Multi-task Masked Autoencoders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E. Hinton · 2020
Cited alongside, same era.
Hierarchical neural architecture search for deep stereo matching
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, and Zongyuan Ge · 2020
Cited alongside, same era.
Bootstrap Your Own Latent - A New Approach to Self-Supervised Learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Cited alongside, same era.
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
A survey on deep learning techniques for stereo-based depth estimation
Hamid Laga, Laurent Valentin Jospin, Farid Boussaid, and Mohammed Bennamoun · 2020
Cited alongside, same era.
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir · 2022
Closest in time.
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli · 2022
Closest in time.
Deep equilibrium optical flow estimation
Shaojie Bai, Zhengyang Geng, Yash Savani, and J. Zico Kolter · 2022
Closest in time.
BEiT: BERT Pre-Training of Image Transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2022
Closest in time.
Masked Autoencoders are Scalable Vision Learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Closest in time.
FlowFormer: A transformer architecture for optical flow
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li · 2022
Closest in time.
High resolution multi-scale raft (robust vision challenge 2022)
Azin Jahedi, Maximilian Luz, Lukas Mehl, Marc Rivinius, and Andrés Bruhn · 2022
Closest in time.
What to Hide from Your Students: Attention-Guided Masked Image Modeling
Ioannis Kakogeorgiou, Spyros Gidaris, Bill Psomas, Yannis Avrithis, Andrei Bursuc, Konstantinos Karantzalos, and Nikos Komodakis · 2022
Closest in time.
Practical stereo matching via cascaded recurrent network with adaptive correlation
Jiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai, Ziwei Yan, Lei Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu · 2022
Closest in time.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al · 2022
Closest in time.
Real-world robot learning with masked visual pre-training
Ilija Radosavovic, Tete Xiao, Stephen James, Pieter Abbeel, Jitendra Malik, and Trevor Darrell · 2022
Closest in time.
Open challenges in deep stereo: the booster dataset
Pierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano · 2022
Closest in time.
Pcw-net: Pyramid combination and warping cost volume for stereo matching
Zhelun Shen, Yuchao Dai21, Xibin Song11, Zhibo Rao, Dingfu Zhou, and Liangjun Zhang · 2022
Closest in time.
Craft: Cross-attentional flow transformer for robust optical flow
Xiuchao Sui, Shaohua Li, Xue Geng, Yan Wu, Xinxing Xu, Yong Liu, Rick Goh, and Hongyuan Zhu · 2022
Closest in time.
SKFlow: Learning optical flow with super kernels
Shangkun Sun, Yuanqi Chen, Yu Zhu, Guodong Guo, and Ge Li · 2022
Closest in time.
Masked Feature Prediction for Self-Supervised Visual Pre-Training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer · 2022
Closest in time.
CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion
Weinzaepfel, Philippe and Leroy, Vincent and Lucas, Thomas and Brégier, Romain and Cabon, Yohann and Arora, Vaibhav and Antsfeld, Leonid and Chidlovskii, Boris and Csurka, Gabriela and Revaud Jérôme · 2022
Closest in time.
SimMIM: A Simple Framework for Masked Image Modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Closest in time.
Attention concatenation volume for accurate and efficient stereo matching
Gangwei Xu, Junda Cheng, Peng Guo, and Xin Yang · 2022
Closest in time.
Gmflow: Learning optical flow via global matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao · 2022
Closest in time.
Vitpose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao · 2022
Closest in time.
Global matching with overlapping attention for optical flow estimation
Shiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou, and Dimitris N. Metaxas · 2022
Closest in time.
DIP: deep inverse patchmatch for high-resolution optical flow
Zihua Zheng, Ni Nie, Zhi Ling, Pengfei Xiong, Jiangyu Liu, Hao Wang, and Jiankun Li · 2022
Closest in time.
iBOT: Image BERT Pre-training with Online Tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong · 2022
Closest in time.
The hidden uniform cluster prior in self-supervised learning
Mahmoud Assran, Randall Balestriero, Quentin Duval, Florian Bordes, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, and Nicolas Ballas · 2023
Closest in time.
Corrupted Image Modeling for Self-Supervised Visual Pre-Training
Yuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang, and Furu Wei · 2023
Closest in time.
Exploring the role of mean teachers in self-supervised masked auto-encoders
Youngwan Lee, Jeffrey Willette, Jonghee Kim, Juho Lee, and Sung Ju Hwang · 2023
Closest in time.
Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo
Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, and Andrés Bruhn · 2023
Closest in time.
Unifying flow, stereo and depth estimation
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger · 2023
Closest in time.