Fetching the paper…
Reading the bibliography…
We present Pow3r, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts.
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
Martin A. Fischler and Robert C. Bolles · 1981
Earlier work this paper cites.
Interpreting perspective images
Stephen T. Barnard · 1983
Earlier work this paper cites.
A combined corner and edge detector
Chris Harris, Mike Stephens, et al · 1988
Earlier work this paper cites.
Procrustes alignment with the em algorithm
Bin Luo and Edwin R. Hancock · 1999
Earlier work this paper cites.
Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman · 2003
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G Lowe · 2004
Earlier work this paper cites.
Surf: Speeded up robust features
Herbert Bay · 2006
Earlier work this paper cites.
Machine learning for high-speed corner detection
Edward Rosten and Tom Drummond · 2006
Earlier work this paper cites.
Epnp: An accurate o(n) solution to the pnp problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua · 2009
Earlier work this paper cites.
Lsd: A fast line segment detector with a false detection control
Rafael Grompone von Gioi, Jérémie Jakubowicz, Jean-Michel Morel, and Gregory Randall · 2010
Earlier work this paper cites.
The weiszfeld algorithm: Proof, amendments, and extensions
Frank Plastria · 2011
Earlier work this paper cites.
Sfm with mrfs: Discrete-continuous optimization for large-scale structure from motion
David J Crandall, Andrew Owens, Noah Snavely, and Daniel P Huttenlocher · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
A global linear method for camera pose registration
Nianjuan Jiang, Zhaopeng Cui, and Ping Tan · 2013
Earlier work this paper cites.
Direct linear transformation from comparator coordinates into object space coordinates in close range photogrammetry
Yousset I Abdel-Aziz, Hauck Michael Karara, and Michael Hauck · 2015
Earlier work this paper cites.
Multi-view stereo: A tutorial
Yasutaka Furukawa, Carlos Hernández, et al · 2015
Earlier work this paper cites.
Massively parallel multiview stereopsis by surface normal diffusion
Silvano Galliani, Katrin Lasinger, and Konrad Schindler · 2015
Earlier work this paper cites.
Large-scale data for multiple-view stereopsis
Henrik Aanæs, Rasmus Ramsbøl Jensen, George Vogiatzis, Engin Tola, and Anders Bjorholm Dahl · 2016
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes L Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys · 2016
Earlier work this paper cites.
Lift: Learned invariant feature transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua · 2016
Earlier work this paper cites.
Detecting vanishing points using global image context in a non-manhattan world
Menghua Zhai, Scott Workman, and Nathan Jacobs · 2016
Earlier work this paper cites.
Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors
Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krystian Mikolajczyk · 2017
Earlier work this paper cites.
Hsfm: Hybrid structure-from-motion
Hainan Cui, Xiang Gao, Shuhan Shen, and Zhanyi Hu · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun · 2017
Earlier work this paper cites.
ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras
Raúl Mur-Artal and Juan D. Tardós · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Demon: Depth and motion network for learning monocular stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox · 2017
Earlier work this paper cites.
Depth estimation via affinity learned with convolutional spatial propagation network
Xinjing Cheng, Peng Wang, and Ruigang Yang · 2018
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Sparse-to-dense: Depth prediction from sparse depth samples and a single image
Fangchang Ma, Guilherme Venturelli Cavalheiro, and Sertac Karaman · 2018
Earlier work this paper cites.
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan · 2018
Earlier work this paper cites.
Deep depth completion of a single rgb-d image
Yinda Zhang and Thomas Funkhouser · 2018
Earlier work this paper cites.
Key. net: Keypoint detection by handcrafted and learned cnn filters
Axel Barroso-Laguna, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk · 2019
Earlier work this paper cites.
Deep single image camera calibration with radial distortion
Manuel López-Antequera, Roger Marí, Pau Gargallo, Yubin Kuang, Javier Gonzalez-Jimenez, and Gloria Haro · 2019
Earlier work this paper cites.
Self-supervised sparse-to-dense: Self-supervised depth completion from lidar and monocular camera
Fangchang Ma, Guilherme Venturelli Cavalheiro, and Sertac Karaman · 2019
Earlier work this paper cites.
Deeplidar: Deep surface normal guided depth prediction for outdoor scene from sparse lidar data and single color image
Jiaxiong Qiu, Zhaopeng Cui, Yinda Zhang, Xingdi Zhang, Shuaicheng Liu, Bing Zeng, and Marc Pollefeys · 2019
Cited alongside, same era.
From coarse to fine: Robust hierarchical localization at large scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk · 2019
Cited alongside, same era.
Habitat: A platform for embodied ai research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al · 2019
Cited alongside, same era.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein · 2019
Cited alongside, same era.
Sparse and noisy lidar completion with rgb guidance and uncertainty
Wouter Van Gansbeke, Davy Neven, Bert De Brabandere, and Luc Van Gool · 2019
Cited alongside, same era.
Rethinking depth estimation for multi-view stereo: A unified representation
Rui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai, and Ronggang Wang · 2022
Later among the works it cites.
Guideformer: Transformers for image guided depth completion
Kyeongha Rho, Jinsung Ha, and Youngjung Kim · 2022
Later among the works it cites.
The 8-point algorithm as an inductive bias for relative pose prediction by vits
Chris Rockwell, Justin Johnson, and David F Fouhey · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
A benchmark and a baseline for robust multi-view depth estimation
Philipp Schröppel, Jan Bechtold, Artemij Amiranashvili, and Thomas Brox · 2022
Later among the works it cites.
Quadtree attention for vision transformers
Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cascade cost volume for high-resolution multi-view stereo and stereo matching
Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan · 2020
Cited alongside, same era.
Consac: Robust multi-model fitting by conditional sample consensus
Florian Kluger, Eric Brachmann, Hanno Ackermann, Carsten Rother, Michael Ying Yang, and Bodo Rosenhahn · 2020
Cited alongside, same era.
Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger · 2020
Cited alongside, same era.
Superglue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2020
Cited alongside, same era.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al · 2020
Cited alongside, same era.
Learning guided convolutional network for depth completion
Jie Tang, Fei-Peng Tian, Wei Feng, Jian Li, and Ping Tan · 2020
Cited alongside, same era.
Deepv2d: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng · 2020
Cited alongside, same era.
Later among the works it cites.
Transformer based line segment classifier with image context for real-time vanishing point detection in manhattan world
Xin Tong, Xianghua Ying, Yongjie Shi, Ruibin Wang, and Jinfa Yang · 2022
Later among the works it cites.
Rignet++: Semantic assisted repetitive image guided network for depth completion
Zhiqiang Yan, Kun Wang, Xiang Li, Zhenyu Zhang, Baobei Xu, Jun Li, and Jian Yang · 2022
Later among the works it cites.
Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions
Zhenpei Yang, Zhile Ren, Qi Shan, and Qixing Huang · 2022
Later among the works it cites.
Towards accurate reconstruction of 3d scene shape from a single monocular image
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Simon Chen, Yifan Liu, and Chunhua Shen · 2022
Later among the works it cites.
Relpose: Predicting probabilistic relative rotation for single objects in the wild
Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani · 2022
Later among the works it cites.
Affineglue: Joint matching and robust estimation
Daniel Barath, Dmytro Mishkin, Luca Cavalli, Paul-Edouard Sarlin, Petr Hruby, and Marc Pollefeys · 2023
Later among the works it cites.
Dbarf: Deep bundle-adjusting generalizable neural radiance fields
Yu Chen and Gim Hee Lee · 2023
Later among the works it cites.
Explicit correspondence matching for generalizable neural radiance fields
Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai · 2023
Later among the works it cites.
Completionformer: Efficient and accurate depth completion via hybrid attention with transformers
Shengyu Jiang, Chen Wang, Yuhong Zhang, and Zhibin Wang · 2023
Later among the works it cites.
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis · 2023
Later among the works it cites.
Vision transformer for nerf-based view synthesis from a single input image
Kai-En Lin, Yen-Chen Lin, Wei-Sheng Lai, Tsung-Yi Lin, Yi-Chang Shih, and Ravi Ramamoorthi · 2023
Later among the works it cites.
Lightglue: Local feature matching at light speed
Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys · 2023
Later among the works it cites.
Camp: Camera preconditioning for neural radiance fields
Keunhong Park, Philipp Henzler, Ben Mildenhall, Jonathan T. Barron, and Ricardo Martin-Brualla · 2023
Later among the works it cites.
Infinite photorealistic worlds using procedural generation
Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, Alejandro Newell, Hei Law, Ankit Goyal, Kaiyu Yang, and Jia Deng · 2023
Later among the works it cites.
Vq3d: Learning a 3d-aware generative model on imagenet
Kyle Sargent, Jing Yu Koh, Han Zhang, Huiwen Chang, Charles Herrmann, Pratul Srinivasan, Jiajun Wu, and Deqing Sun · 2023
Later among the works it cites.
Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow
Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, and Jérôme Revaud · 2023
Later among the works it cites.
Scannet++: A high-fidelity dataset of 3d indoor scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai · 2023
Later among the works it cites.
Tame a wild camera: In-the-wild monocular camera calibration
Shengjie Zhu, Abhinav Kumar, Masa Hu, and Xiaoming Liu · 2023
Later among the works it cites.
Learning structure-from-motion with graph attention networks
Lucas Brynte, José Pedro Iglesias, Carl Olsson, and Fredrik Kahl · 2024
Later among the works it cites.
Mvsformer++: Revealing the devil in transformer’s details for multi-view stereo
Chenjie Cao, Xinlin Ren, and Yanwei Fu · 2024
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Timothée Darcet, Mathilde Caron, Piotr Bojanowski, Ishan Misra, Armand Joulin, Julien Mairal, Hervé Jégou, and Hugo Touvron · 2024
Later among the works it cites.
Nvist: In the wild new view synthesis from a single image with transformers
Wonbong Jang and Lourdes Agapito · 2024
Later among the works it cites.
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani · 2024
Later among the works it cites.
Chatgpt-4, 2024
OpenAI · 2024
Later among the works it cites.
Unidepth: Universal monocular metric depth estimation
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu · 2024
Later among the works it cites.
Splatt3r: Zero-shot gaussian splatting from uncalibrated image pairs
Brandon Smart, Chuanxia Zheng, Iro Laina, and Victor Adrian Prisacariu · 2024
Later among the works it cites.
Bilateral propagation network for depth completion
Jie Tang, Fei-Peng Tian, Boshi An, Jian Li, and Ping Tan · 2024
Later among the works it cites.
GeoCalib: Single-image Calibration with Geometric Optimization
Alexander Veicht, Paul-Edouard Sarlin, Philipp Lindenberger, and Marc Pollefeys · 2024
Later among the works it cites.
Tri-perspective view decomposition for geometry-aware depth completion
Zhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng, Yufei Wang, Zhenyu Zhang, Jun Li, and Jian Yang · 2024
Later among the works it cites.
Cameras as rays: Pose estimation via ray diffusion
Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani · 2024
Later among the works it cites.
MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud · 2025
Closest in time.
3d reconstruction with spatial memory
Hengyi Wang and Lourdes Agapito · 2025
Closest in time.
Monst3r: A simple approach for estimating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang · 2025
Closest in time.