Fetching the paper…
Reading the bibliography…
We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos.
In defense of the eight-point algorithm
Richard I Hartley · 1997
Earlier work this paper cites.
Object Recognition from Local Scale-Invariant Features
David G. Lowe · 1999
Earlier work this paper cites.
Bundle adjustment—a modern synthesis
Bill Triggs, Philip F McLauchlan, Richard I Hartley, and Andrew W Fitzgibbon · 2000
Earlier work this paper cites.
Multiple View Geometry in Computer Vision
Richard Hartley and Andrew Zisserman · 2004
Earlier work this paper cites.
Distinctive Image Features from Scale-Invariant Keypoints
David G. Lowe · 2004
Earlier work this paper cites.
An efficient solution to the five-point relative pose problem
David Nistér · 2004
Earlier work this paper cites.
Five-point motion estimation made easy
Hongdong Li and Richard Hartley · 2006
Earlier work this paper cites.
Speeded-Up Robust Features (SURF)
Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool · 2008
Earlier work this paper cites.
Particle video: Long-range motion estimation using point trajectories
Peter Sand and Seth Teller · 2008
Earlier work this paper cites.
Sift flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba · 2010
Earlier work this paper cites.
A benchmark for the evaluation of RGB-D SLAM systems
Jürgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Ssd-6d: Making rgb-based 3d detection and 6d pose estimation great again
Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab · 2017
Earlier work this paper cites.
A survey of structure from motion*
Onur Özyeşil, Vladislav Voroninski, Ronen Basri, and Amit Singer · 2017
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J. Black · 2017
Earlier work this paper cites.
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox · 2017
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Earlier work this paper cites.
Deepmvs: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang · 2018
Earlier work this paper cites.
ReFusion: 3D Reconstruction in Dynamic Environments for RGB-D Cameras Exploiting Residuals
E. Palazzolo, J. Behley, P. Lottes, P. Giguère, and C. Stachniss · 2019
Earlier work this paper cites.
Virtual kitti 2, 2020
Yohann Cabon, Naila Murray, and Martin Humenberger · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang · 2021
Cited alongside, same era.
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotný · 2021
Cited alongside, same era.
Tap-vid: A benchmark for tracking any point in a video
Bootstap: Bootstrapped training for tracking-any-point
Carl Doersch, Pauline Luc, Yi Yang, Dilara Gokay, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ignacio Rocco, Ross Goroshin, João Carreira, et al · 2024
Later among the works it cites.
Track4gen: Teaching video diffusion models to track points improves video generation
Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye, Niloy J. Mitra, and Duygu Ceylan · 2024
Later among the works it cites.
Stereo4d: Learning how things move in 3d from internet stereo videos
Linyi Jin, Richard Tucker, Zhengqi Li, David Fouhey, Noah Snavely, and Aleksander Holynski · 2024
Later among the works it cites.
Fast encoder-based 3d from casual videos via point track processing
Yoni Kasten, Wuyue Lu, and Haggai Maron · 2024
Later among the works it cites.
Tapvid-3d: A benchmark for tracking any point in 3d
Skanda Koppula, Ignacio Rocco, Yi Yang, Joseph Heyward, João Carreira, Andrew Zisserman, Gabriel Brostow, and Carl Doersch · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Recasens, Lucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Cited alongside, same era.
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al · 2022
Cited alongside, same era.
Particle video revisited: Tracking through occlusions using point trajectories
Adam W. Harley, Zhaoyuan Fang, and Katerina Fragkiadaki · 2022
Cited alongside, same era.
HOI4D: A 4d egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi · 2022
Cited alongside, same era.
Virtual correspondence: Humans as a cue for extreme-view geometry
Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, and Antonio Torralba · 2022
Cited alongside, same era.
Neural window fully-connected crfs for monocular depth estimation
Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and Ping Tan · 2022
Cited alongside, same era.
Later among the works it cites.
Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds
Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis · 2024
Later among the works it cites.
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al · 2024
Later among the works it cites.
Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation
Haozhe Lou, Yurong Liu, Yike Pan, Yiran Geng, Jianteng Chen, Wenlong Ma, Chenglong Li, Lin Wang, Hengzhen Feng, Lu Shi, Liyi Luo, and Yongliang Shi · 2024
Later among the works it cites.
Delta: Dense efficient long-range 3d tracking for any video
Tuan Duc Ngo, Peiye Zhuang, Chuang Gan, Evangelos Kalogerakis, Sergey Tulyakov, Hsin-Ying Lee, and Chaoyang Wang · 2024
Later among the works it cites.
Codef: Content deformation fields for temporally consistent video processing
Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Juntao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen · 2024
Later among the works it cites.
Unidepth: Universal monocular metric depth estimation
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segù, Siyuan Li, Luc Van Gool, and Fisher Yu · 2024
Later among the works it cites.
Kalib: Markerless hand-eye calibration with keypoint tracking
Tutian Tang, Minghao Liu, Wenqiang Xu, and Cewu Lu · 2024
Later among the works it cites.
Marigold-dc: Zero-shot monocular depth completion with guided diffusion
Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, and Anton Obukhov · 2024
Later among the works it cites.
RGBD objects in the wild: Scaling real-world 3d object learning from RGB-D videos
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang · 2024
Later among the works it cites.
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou · 2024
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
L4p: Low-level 4d vision perception unified
Abhishek Badki, Hang Su, Bowen Wen, and Orazio Gallo · 2025
Closest in time.
Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yijin Li, Fu-Yun Wang, and Hongsheng Li · 2025
Closest in time.
Video depth anything: Consistent depth estimation for super-long videos
Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, and Bingyi Kang · 2025
Closest in time.
Diffusion as shader: 3d-aware video diffusion for versatile video generation control
Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu · 2025
Closest in time.
Megasam: Accurate, fast and robust structure and motion from casual dynamic videos
Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holynski, and Noah Snavely · 2025
Closest in time.
Pre-training auto-regressive robotic models with 4d representations
Dantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby, Junyi Zhang, Ziteng Ji, Trevor Darrell, and Roei Herzig · 2025
Closest in time.
Dynamic camera poses and where to find them
Chris Rockwell, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu, David F. Fouhey, and Chen-Hsuan Lin · 2025
Closest in time.
Depth anything v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2025
Closest in time.
Tapip3d: Tracking any point in persistent 3d geometry
Bowei Zhang, Lei Ke, Adam W Harley, and Katerina Fragkiadaki · 2025
Closest in time.