Fetching the paper…
Reading the bibliography…
Feature representation learning is the key recipe for learning-based Multi-View Stereo (MVS).
Patchmatch: A randomized correspondence algorithm for structural image editing
Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman · 2009
Earlier work this paper cites.
Accurate, dense, and robust multiview stereopsis
Yasutaka Furukawa and Jean Ponce · 2009
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Multi-view stereo: A tutorial
Yasutaka Furukawa and Carlos Hernández · 2015
Earlier work this paper cites.
Massively parallel multiview stereopsis by surface normal diffusion
Silvano Galliani, Katrin Lasinger, and Konrad Schindler · 2015
Earlier work this paper cites.
Large-scale data for multiple-view stereopsis
Henrik Aanæs, Rasmus Ramsbøl Jensen, George Vogiatzis, Engin Tola, and Anders Bjorholm Dahl · 2016
Earlier work this paper cites.
A unified multi-scale deep convolutional neural network for fast object detection
Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes L Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys · 2016
Earlier work this paper cites.
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Earlier work this paper cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Earlier work this paper cites.
Surfacenet: An end-to-end 3d neural network for multiview stereopsis
Mengqi Ji, Juergen Gall, Haitian Zheng, Yebin Liu, and Lu Fang · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
End-to-end learning of geometry and context for deep stereo regression
Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Peter Henry, Ryan Kennedy, Abraham Bachrach, and Adam Bry · 2017
Earlier work this paper cites.
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
Yolo9000: better, faster, stronger
Joseph Redmon and Ali Farhadi · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan · 2018
Earlier work this paper cites.
Group-wise correlation stereo network
Xiaoyang Guo, Kai Yang, Wukui Yang, Xiaogang Wang, and Hongsheng Li · 2019
Earlier work this paper cites.
P-mvsnet: Learning patch-wise matching confidence aggregation for multi-view stereo
Keyang Luo, Tao Guan, Lili Ju, Haipeng Huang, and Yawei Luo · 2019
Earlier work this paper cites.
Multi-scale geometric consistency guided multi-view stereo
Qingshan Xu and Wenbing Tao · 2019
Cited alongside, same era.
Recurrent mvsnet for high-resolution multi-view stereo depth inference
Yao Yao, Zixin Luo, Shiwei Li, Tianwei Shen, Tian Fang, and Long Quan · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Pct: Point cloud transformer
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Cotr: Correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi · 2021
Later among the works it cites.
Patchmatch-rl: Deep mvs with pixelwise depth, normal, and visibility
Jae Yong Lee, Joseph DeGol, Chuhang Zou, and Derek Hoiem · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Cascade cost volume for high-resolution multi-view stereo and stereo matching
Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan · 2020
Cited alongside, same era.
How much position information do convolutional neural networks encode?
Md Amirul Islam, Sen Jia, and Neil DB Bruce · 2020
Cited alongside, same era.
Glu variants improve transformer
Noam Shazeer · 2020
Cited alongside, same era.
Bp-mvsnet: Belief-propagation-layers for multi-view-stereo
Christian Sormann, Patrick Knöbelreiter, Andreas Kuhn, Mattia Rossi, Thomas Pock, and Friedrich Fraundorfer · 2020
Cited alongside, same era.
Pvsnet: Pixelwise visibility-aware multi-view stereo network
Qingshan Xu and Wenbing Tao · 2020
Cited alongside, same era.
Marmvs: Matching ambiguity reduced multiple view stereo for efficient large scale scene reconstruction
Zhenyu Xu, Yiguang Liu, Xuelei Shi, Ying Wang, and Yunan Zheng · 2020
Cited alongside, same era.
Xinjun Ma, Yue Gong, Qirui Wang, Jingwei Huang, Lei Chen, and Fan Yu · 2021
Later among the works it cites.
Generalized binary search network for highly-efficient multi-view stereo
Zhenxing Mi, Di Chang, and Dan Xu · 2021
Later among the works it cites.
Loftr: Detector-free local feature matching with transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou · 2021
Later among the works it cites.
Aa-rmvsnet: Adaptive aggregation recurrent multi-view stereo network
Zizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen, and Guoping Wang · 2021
Later among the works it cites.
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao · 2021
Later among the works it cites.
Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions
Zhenpei Yang, Zhile Ren, Qi Shan, and Qixing Huang · 2021
Later among the works it cites.
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2021
Later among the works it cites.
Multi-view stereo with transformer
Jie Zhu, Bo Peng, Wanqing Li, Haifeng Shen, Zhe Zhang, and Jianjun Lei · 2021
Later among the works it cites.
Learning prior feature and attention enhanced image inpainting
Chenjie Cao, Qiaole Dong, and Yanwei Fu · 2022
Closest in time.
Any-resolution training for high-resolution image synthesis
Lucy Chai, Michael Gharbi, Eli Shechtman, Phillip Isola, and Richard Zhang · 2022
Closest in time.
Kd-mvs: Knowledge distillation based self-supervised learning for mvs
Yikang Ding, Qingtian Zhu, Xiangyue Liu, Wentao Yuan, Haotian Zhang, and CHi Zhang · 2022
Closest in time.
Incremental transformer structure enhanced image inpainting with masking positional encoding
Qiaole Dong, Chenjie Cao, and Yanwei Fu · 2022
Closest in time.
Flowformer: A transformer architecture for optical flow
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li · 2022
Closest in time.
Exploring plain vision transformer backbones for object detection
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He · 2022
Closest in time.
Rethinking depth estimation for multi-view stereo: A unified representation and focal loss
Rui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai, and Ronggang Wang · 2022
Closest in time.
Bidirectional hybrid lstm based recurrent neural network for multi-view stereo
Zizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen, and Guoping Wang · 2022
Closest in time.