Fetching the paper…
Reading the bibliography…
We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views.
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
Martin A Fischler and Robert C Bolles · 1981
Earlier work this paper cites.
Least-squares estimation of transformation parameters between two point patterns
Shinji Umeyama · 1991
Earlier work this paper cites.
Multiple View Geometry in Computer Vision
Richard Hartley and Andrew Zisserman · 2000
Earlier work this paper cites.
A critique of structure-from-motion algorithms
John Oliensis · 2000
Earlier work this paper cites.
Linear multiview reconstruction of points, lines, planes and cameras using a reference plane
Rother · 2003
Earlier work this paper cites.
Multiple View Geometry in Computer Vision
Richard Hartley and Andrew Zisserman · 2004
Earlier work this paper cites.
Photo tourism: exploring photo collections in 3d
Noah Snavely, Steven M Seitz, and Richard Szeliski · 2006
Earlier work this paper cites.
Particle video: Long-range motion estimation using point trajectories
Peter Sand and Seth Teller · 2008
Earlier work this paper cites.
Ep n p: An accurate o (n) solution to the p n p problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua · 2009
Earlier work this paper cites.
Building rome on a cloudless day
Jan-Michael Frahm, Pierre Fite-Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik, et al · 2010
Earlier work this paper cites.
Building rome in a day
Sameer Agarwal, Yasutaka Furukawa, Noah Snavely, Ian Simon, Brian Curless, Steven M Seitz, and Richard Szeliski · 2011
Earlier work this paper cites.
Global motion estimation from point matches
Mica Arie-Nachimson, Shahar Z Kovalsky, Ira Kemelmacher-Shlizerman, Amit Singer, and Ronen Basri · 2012
Earlier work this paper cites.
Sfm with mrfs: Discrete-continuous optimization for large-scale structure from motion
David J Crandall, Andrew Owens, Noah Snavely, and Daniel P Huttenlocher · 2012
Earlier work this paper cites.
A global linear method for camera pose registration
Nianjuan Jiang, Zhaopeng Cui, and Ping Tan · 2013
Earlier work this paper cites.
Global fusion of relative motions for robust, accurate and scalable structure from motion
Pierre Moulon, Pascal Monasse, and Renaud Marlet · 2013
Earlier work this paper cites.
Towards linear-time incremental structure from motion
Changchang Wu · 2013
Earlier work this paper cites.
Large scale multi-view stereopsis evaluation
Rasmus Jensen, Anders Dahl, George Vogiatzis, Engil Tola, and Henrik Aanæs · 2014
Earlier work this paper cites.
Global structure-from-motion by similarity averaging
Zhaopeng Cui and Ping Tan · 2015
Earlier work this paper cites.
Linear global translation estimation with feature tracks
Zhaopeng Cui, Nianjuan Jiang, Chengzhou Tang, and Ping Tan · 2015
Earlier work this paper cites.
Multi-view stereo: A tutorial
Yasutaka Furukawa, Carlos Hernández, et al · 2015
Earlier work this paper cites.
Robust camera location estimation by convex programming
Onur Ozyesil and Amit Singer · 2015
Earlier work this paper cites.
Optimizing the viewing graph for structure-from-motion
Chris Sweeney, Torsten Sattler, Tobias Hollerer, Matthew Turk, and Marc Pollefeys · 2015
Earlier work this paper cites.
Modelling uncertainty in deep learning for camera relocalization
Alex Kendall and Roberto Cipolla · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes L Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys · 2016
Earlier work this paper cites.
LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua · 2016
Earlier work this paper cites.
Hsfm: Hybrid structure-from-motion
Hainan Cui, Xiang Gao, Shuhan Shen, and Zhanyi Hu · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
What uncertainties do we need in Bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
Learning 3D object categories by looking around them
David Novotný, Diane Larlus, and Andrea Vedaldi · 2017
Earlier work this paper cites.
A survey of structure from motion*
Onur Özyeşil, Vladislav Voroninski, Ronen Basri, and Amit Singer · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Demon: Depth and motion network for learning monocular stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox · 2017
Earlier work this paper cites.
Unsupervised learning of depth and ego-motion from video
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe · 2017
Earlier work this paper cites.
Superpoint: Self-supervised interest point detection and description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Deepmvs: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang · 2018
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Capturing the geometry of object categories from video supervision
David Novotný, Diane Larlus, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Ba-net: Dense bundle adjustment network
Chengzhou Tang and Ping Tan · 2018
Earlier work this paper cites.
Deepv2d: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
D2-net: A trainable cnn for joint description and detection of local features
Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler · 2019
Cited alongside, same era.
The replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown · 2020
Cited alongside, same era.
Yohann Cabon, Naila Murray, and Martin Humenberger · 2020
Cited alongside, same era.
Cascade cost volume for high-resolution multi-view stereo and stereo matching
Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan · 2020
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Improving transformer-based image matching by cascaded capturing spatially informative keypoints
Chenjie Cao and Yanwei Fu · 2023
Later among the works it cites.
Vision transformers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2023
Later among the works it cites.
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi · 2023
Later among the works it cites.
DKM: Dense kernelized feature matching for geometry estimation
Johan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck, and Michael Felsberg · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Query-key normalization for transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen · 2020
Cited alongside, same era.
Mapillary planet-scale depth dataset
Manuel Lopez-Antequera, Pau Gargallo, Markus Hofinger, Samuel Rota Bulò, Yubin Kuang, and Peter Kontschieder · 2020
Cited alongside, same era.
Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger · 2020
Cited alongside, same era.
Disk: Learning local features with policy gradient
Michał Tyszkiewicz, Pascal Fua, and Eduard Trulls · 2020
Cited alongside, same era.
Deepsfm: Structure from motion via deep bundle adjustment
Xingkui Wei, Yinda Zhang, Zhuwen Li, Yanwei Fu, and Xiangyang Xue · 2020
Cited alongside, same era.
Learning inverse depth regression for multi-view stereo with correlation cost volume
Qingshan Xu and Wenbing Tao · 2020
Cited alongside, same era.
Blendedmvs: A large-scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan · 2020
Cited alongside, same era.
Later among the works it cites.
Detector-free structure from motion
Xingyi He, Jiaming Sun, Yifan Wang, Sida Peng, Qixing Huang, Hujun Bao, and Xiaowei Zhou · 2023
Later among the works it cites.
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani · 2023
Later among the works it cites.
Aria digital twin: A new benchmark dataset for egocentric 3d machine perception
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Yuheng (Carl) Ren · 2023
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
Sparsepose: Sparse-view camera pose regression and refinement
Samarth Sinha, Jason Y Zhang, Andrea Tagliasacchi, Igor Gilitschenski, and David B Lindell · 2023
Later among the works it cites.
Gokul Yenduri, Ramalingam M, Chemmalar Selvi G., Supriya Y, Gautam Srivastava, Praveen Kumar Reddy Maddikunta, Deepti Raj G, Rutvij H. Jhaveri, Prabadevi B, Weizheng Wang, Athanasios V. Vasilakos, and Thippa Reddy Gadekallu · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Joshua M Susskind · 2023
Later among the works it cites.
Geomvsnet: Learning multi-view stereo with geometry perception
Zhe Zhang, Rui Peng, Yuxi Hu, and Ronggang Wang · 2023
Later among the works it cites.
Aliked: A lighter keypoint and descriptor extraction network via deformable transformation
Xiaoming Zhao, Xingming Wu, Weihai Chen, Peter CY Chen, Qingsong Xu, and Zhengguo Li · 2023
Later among the works it cites.
Pointodyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J. Guibas · 2023
Later among the works it cites.
Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer
Eric Brachmann, Jamie Wynn, Shuai Chen, Tommaso Cavallari, Áron Monszpart, Daniyar Turmukhambetov, and Victor Adrian Prisacariu · 2024
Later among the works it cites.
Local all-pair correspondence for point tracking
Seokju Cho, Jiahui Huang, Jisu Nam, Honggyu An, Seungryong Kim, and Joon-Young Lee · 2024
Later among the works it cites.
Bootstap: Bootstrapped training for tracking-any-point
Carl Doersch, Yi Yang, Dilara Gokay, Pauline Luc, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ross Goroshin, João Carreira, and Andrew Zisserman · 2024
Later among the works it cites.
MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud · 2024
Later among the works it cites.
Roma: Robust dense feature matching
Johan Edstedt, Qiyu Sun, Georg Bökman, Mårten Wadenbäck, and Michael Felsberg · 2024
Later among the works it cites.
Flex3d: Feed-forward 3d generation with flexible reconstruction model and input view curation
Junlin Han, Jianyuan Wang, Andrea Vedaldi, Philip Torr, and Filippos Kokkinos · 2024
Later among the works it cites.
LRM: Large reconstruction model for single image to 3D
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan · 2024
Later among the works it cites.
LVSM: a large view synthesis model with minimal 3D inductive bias
Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu · 2024
Later among the works it cites.
Dense optical tracking: Connecting the dots
Guillaume Le Moing, Jean Ponce, and Cordelia Schmid · 2024
Later among the works it cites.
Grounding image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jérôme Revaud · 2024
Later among the works it cites.
Taptr: Tracking any point with transformers as detection
Hongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, and Lei Zhang · 2024
Later among the works it cites.
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al · 2024
Later among the works it cites.
DINOv2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Later among the works it cites.
Global Structure-from-Motion Revisited
Linfei Pan, Daniel Barath, Marc Pollefeys, and Johannes Lutz Schönberger · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Later among the works it cites.
Splatter image: Ultra-fast single-view 3d reconstruction
Stanislaw Szymanowicz, Chrisitian Rupprecht, and Andrea Vedaldi · 2024
Later among the works it cites.
MeshLRM: large reconstruction model for high-quality mesh
Xinyue Wei, Kai Zhang, Sai Bi, Hao Tan, Fujun Luan, Valentin Deschaintre, Kalyan Sunkavalli, Hao Su, and Zexiang Xu · 2024
Later among the works it cites.
Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang · 2024
Later among the works it cites.
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou · 2024
Later among the works it cites.
GRM: Large gaussian reconstruction model for efficient 3D reconstruction and generation
Yinghao Xu, Zifan Shi, Wang Yifan, Hansheng Chen, Ceyuan Yang, Sida Peng, Yujun Shen, and Gordon Wetzstein · 2024
Later among the works it cites.
Lightplane: Highly-scalable components for neural 3Dfields
Ang Cao, Justin Johnson, Andrea Vedaldi, and David Novotny · 2025
Closest in time.
Robust incremental structure-from-motion with hybrid features
Shaohui Liu, Yidan Gao, Tianyi Zhang, Rémi Pautrat, Johannes L Schönberger, Viktor Larsson, and Marc Pollefeys · 2025
Closest in time.
Continuous 3d perception model with persistent state, 2025
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, and Angjoo Kanazawa · 2025
Closest in time.
Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass
Jianing Yang, Alexander Sax, Kevin J Liang, Mikael Henaff, Hao Tang, Ang Cao, Joyce Chai, Franziska Meier, and Matt Feiszli · 2025
Closest in time.
Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views, 2025
Shangzhan Zhang, Jianyuan Wang, Yinghao Xu, Nan Xue, Christian Rupprecht, Xiaowei Zhou, Yujun Shen, and Gordon Wetzstein · 2025
Closest in time.