Fetching the paper…
Reading the bibliography…
We introduce MapAnything, a unified transformer-based feed-forward model that ingests one or more images along with optional geometric inputs such as camera intrinsics, poses, depth, or partial reconstructions, and then directly regresses the metric 3D scene geometry and cameras.
Photometric method for determining surface orientation from multiple images
Robert J. Woodham · 1980
Earlier work this paper cites.
Obtaining shape from shading information
Berthold KP Horn · 1989
Earlier work this paper cites.
Bundle adjustment – a modern synthesis
Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew Fitzgibbon · 2000
Earlier work this paper cites.
A general imaging model and a method for finding its parameters
Michael D. Grossberg and Shree K. Nayar · 2001
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G. Lowe · 2004
Earlier work this paper cites.
An efficient solution to the five-point relative pose problem
David Nistér · 2004
Earlier work this paper cites.
Geometric context from a single image
Derek Hoiem, Alexei A Efros, and Martial Hebert · 2005
Earlier work this paper cites.
Fast image-based localization using direct 2D-to-3D matching
Torsten Sattler, Bastian Leibe, and Leif Kobbelt · 2011
Earlier work this paper cites.
Rotation averaging
Richard Hartley, Jochen Trumpf, Yuchao Dai, and Hongdong Li · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L. Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes L. Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys · 2016
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schöps, Johannes L. Schönberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
DeMoN: Depth and motion network for learning monocular stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox · 2017
Earlier work this paper cites.
DeepMVS: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang · 2018
Earlier work this paper cites.
MegaDepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Amir R. Zamir, Alexander Sax, William Shen, Leonidas J. Guibas, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
A general and adaptive robust loss function
Jonathan T. Barron · 2019
Earlier work this paper cites.
Mapillary planet-scale depth dataset
Manuel López Antequera, Pau Gargallo, Markus Hofinger, Samuel Rota Bulò, Yubin Kuang, and Peter Kontschieder · 2020
Earlier work this paper cites.
SuperGlue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2020
Earlier work this paper cites.
DeepV2D: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
Neural ray surfaces for self-supervised learning of depth and ego-motion
Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon · 2020
Earlier work this paper cites.
TartanAir: A dataset to push the limits of visual SLAM
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer · 2020
Earlier work this paper cites.
BlendedMVS: A large-scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan · 2020
Earlier work this paper cites.
Fast-MVSNet: Sparse-to-dense multi-view stereo with learned propagation and Gauss-Newton refinement
Zehao Yu and Shenghua Gao · 2020
Earlier work this paper cites.
DeepTAM: Deep tracking and mapping with convolutional neural networks
Huizhong Zhou, Benjamin Ummenhofer, and Thomas Brox · 2020
Cited alongside, same era.
SAIL-VOS 3D: A synthetic dataset and baselines for object detection and 3D mesh reconstruction from video data
Yuan-Ting Hu, Jiahong Wang, Raymond A. Yeh, and Alexander G. Schwing · 2021
Cited alongside, same era.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Cited alongside, same era.
SMD-nets: Stereo mixture density networks
Fabio Tosi, Yiyi Liao, Carolin Schmitt, and Andreas Geiger · 2021
Cited alongside, same era.
MultiMAE: Multi-modal multi-task masked autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir · 2022
Cited alongside, same era.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
GeoCalib: Single-image calibration with geometric optimization
Alexander Veicht, Paul-Edouard Sarlin, Philipp Lindenberger, and Marc Pollefeys · 2024
Later among the works it cites.
Depth anything V2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Cameras as rays: Pose estimation via ray diffusion
Jason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani · 2024
Later among the works it cites.
RayFronts: Open-set semantic ray frontiers for online scene understanding and exploration
Omar Alama, Avigyan Bhattacharya, Haoyang He, Seungchan Kim, Yuheng Qiu, Wenshan Wang, Cherie Ho, Nikhil Keetha, and Sebastian Scherer · 2025
Closest in time.
Depth pro: Sharp monocular metric depth in less than a second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun · 2025
Closest in time.
MUSt3R: Multi-view network for stereo 3D reconstruction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2022
Cited alongside, same era.
A benchmark and a baseline for robust multi-view depth estimation
Philipp Schröppel, Jan Bechtold, Artemij Amiranashvili, and Thomas Brox · 2022
Cited alongside, same era.
Toward general-purpose robots via foundation models: A survey and meta-analysis
Yafei Hu, Quanting Xie, Vidhi Jain, Jonathan Francis, Jay Patrikar, Nikhil Keetha, Seungchan Kim, Yaqi Xie, Tianyi Zhang, Hao-Shu Fang, Shibo Zhao, Shayegan Omidshafiei, Dong-Ki Kim, Ali akbar Agha-mohammadi, Katia Sycara, Matthew Johnson-Roberson, Dhruv Batra, Xiaolong Wang, Sebastian Scherer, Chen Wang, Zsolt Kira, Fei Xia, and Yonatan Bisk · 2023
Cited alongside, same era.
DynamicStereo: Consistent dynamic depth from stereo videos
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2023
Cited alongside, same era.
Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo
Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, and Andrés Bruhn · 2023
Cited alongside, same era.
CroCo v2: Improved cross-view completion pre-training for stereo matching and optical flow
Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, and Jerome Revaud · 2023
Cited alongside, same era.
ScanNet++: A high-fidelity dataset of 3D indoor scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai · 2023
Cited alongside, same era.
Yohann Cabon, Lucas Stoffl, Leonid Antsfeld, Gabriela Csurka, Boris Chidlovskii, Jerome Revaud, and Vincent Leroy · 2025
Closest in time.
Reloc3r: Large-scale training of relative camera pose regression for generalizable, fast, and accurate visual localization
Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, and Yanchao Yang · 2025
Closest in time.
Light3R-SfM: Towards feed-forward structure-from-motion
Sven Elflein, Qunjie Zhou, Sérgio Agostinho, and Laura Leal-Taixé · 2025
Closest in time.
RADIOv2.5: Improved baselines for agglomerative vision foundation models
Greg Heinrich, Mike Ranzinger, Hongxu, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov · 2025
Closest in time.
MVSAnywhere: Zero-shot multi-view stereo
Sergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando, Daniyar Turmukhambetov, Javier Civera, Oisin Mac Aodha, Gabriel Brostow, and Jamie Watson · 2025
Closest in time.
Pow3R: Empowering unconstrained 3D reconstruction with camera and scene priors
Wonbong Jang, Philippe Weinzaepfel, Vincent Leroy, Lourdes Agapito, and Jerome Revaud · 2025
Closest in time.
LVSM: A large view synthesis model with minimal 3D inductive bias
Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu · 2025
Closest in time.
Any4D: Unified feed-forward metric 4d reconstruction
Jay Karhade, Nikhil Keetha, Yuchen Zhang, Tanisha Gupta, Akash Sharma, Sebastian Scherer, and Deva Ramanan · 2025
Closest in time.
Depth anything 3: Recovering the visual space from any views
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang · 2025
Closest in time.
MASt3R-SLAM: Real-time dense SLAM with 3D reconstruction priors
Riku Murai, Eric Dexheimer, and Andrew J. Davison · 2025
Closest in time.
MP-SfM: Monocular surface priors for robust structure-from-motion
Zador Pataki, Paul-Edouard Sarlin, Johannes L. Schönberger, and Marc Pollefeys · 2025
Closest in time.
MV-DUSt3R+: Single-stage scene reconstruction from sparse views in 2 seconds
Zhenggang Tang, Yuchen Fan, Dilin Wang, Hongyu Xu, Rakesh Ranjan, Alexander Schwing, and Zhicheng Yan · 2025
Closest in time.
AnyCalib: On-manifold learning for model-agnostic single-view camera calibration
Javier Tirado-Garín and Javier Civera · 2025
Closest in time.
3D reconstruction with spatial memory
Hengyi Wang and Lourdes Agapito · 2025
Closest in time.
Fast3R: Towards 3D reconstruction of 1000+ images in one forward pass
Jianing Yang, Alexander Sax, Kevin J. Liang, Mikael Henaff, Hao Tang, Ang Cao, Joyce Chai, Franziska Meier, and Matt Feiszli · 2025
Closest in time.
UFM: A simple path towards unified dense correspondence with flow
Yuchen Zhang, Nikhil Keetha, Chenwei Lyu, Bhuvan Jhamb, Yutian Chen, Yuheng Qiu, Jay Karhade, Shreyas Jha, Yaoyu Hu, Deva Ramanan, Sebastian Scherer, and Wenshan Wang · 2025
Closest in time.
Stable virtual camera: Generative view synthesis with diffusion models
Jinghao (Jensen) Zhou, Hang Gao, Vikram Voleti, Aaryaman Vasishta, Chun-Han Yao, Mark Boss, Philip Torr, Christian Rupprecht, and Varun Jampani · 2025
Closest in time.
UniDepthV2: Universal monocular metric depth estimation made simpler
Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool · 2026
Closest in time.
Fillerbuster: Multi-view scene completion for casual captures
Ethan Weber, Norman Müller, Yash Kant, Vasu Agrawal, Michael Zollhöfer, Angjoo Kanazawa, and Christian Richardt · 2026
Closest in time.