Fetching the paper…
Reading the bibliography…
The spatio-temporal relationship between the pixels of a video carries critical information for low-level 4D perception tasks.
A general imaging model and a method for finding its parameters
M.D. Grossberg and S.K. Nayar · 2001
Earlier work this paper cites.
Photo tourism: exploring photo collections in 3d
Noah Snavely, Steven M. Seitz, and Richard Szeliski · 2006
Earlier work this paper cites.
Accurate, dense, and robust multiview stereopsis
Yasutaka Furukawa and Jean Ponce · 2009
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Learning to segment moving objects in videos
Katerina Fragkiadaki, Pablo Arbelaez, Panna Felsen, and Jitendra Malik · 2015
Earlier work this paper cites.
Massively parallel multiview stereopsis by surface normal diffusion
Silvano Galliani, Katrin Lasinger, and Konrad Schindler · 2015
Earlier work this paper cites.
Panoptic Studio: A massively multiview system for social motion capture
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh · 2015
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
N. Mayer, E. Ilg, P. Häusser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes L Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm · 2016
Earlier work this paper cites.
FusionSeg: Learning to combine motion and appearance for fully automatic segmentation of generic objects in videos
Suyog Jain, Bo Xiong, and Kristen Grauman · 2017
Earlier work this paper cites.
Deepmvs: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang · 2018
Earlier work this paper cites.
Unsupervised learning of multi-frame optical flow with occlusions
Joel Janai, Fatma Guney, Anurag Ranjan, Michael Black, and Andreas Geiger · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Earlier work this paper cites.
Deepv2d: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng · 2018
Earlier work this paper cites.
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan · 2018
Earlier work this paper cites.
MoA-Net: Self-supervised motion segmentation
Pia Bideau, Rakesh R. Menon, and Erik Learned-Miller · 2019
Earlier work this paper cites.
A fusion approach for multi-frame optical flow estimation
Zhile Ren, Orazio Gallo, Deqing Sun, Ming-Hsuan Yang, Erik B Sudderth, and Jan Kautz · 2019
Earlier work this paper cites.
Virtual KITTI 2
Yohann Cabon, Naila Murray, and Martin Humenberger · 2020
Earlier work this paper cites.
Consistent video depth estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf · 2020
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2020
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurélien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov · 2020
Earlier work this paper cites.
RAFT: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexander Kolesnikov, Alexey Dosovitskiy, Dirk Weissenborn, Georg Heigold, Jakob Uszkoreit, Lucas Beyer, Matthias Minderer, Mostafa Dehghani, Neil Houlsby, Sylvain Gelly, Thomas Unterthiner, and Xiaohua Zhai · 2021
Cited alongside, same era.
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang · 2021
Cited alongside, same era.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Cited alongside, same era.
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng · 2021
Cited alongside, same era.
Learning to segment rigid motions from two frames
Gengshan Yang and Deva Ramanan · 2021
Cited alongside, same era.
TAP-Vid: A benchmark for tracking any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adrià Recasens, Lucas Smaira, Yusuf Aytar, João Carreira, Andrew Zisserman, and Yi Yang · 2022
Unifying flow, stereo and depth estimation
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger · 2023
Later among the works it cites.
PointOdyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J. Guibas · 2023
Later among the works it cites.
Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation
Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta, and Shubham Tulsiani · 2024
Later among the works it cites.
BootsTAP: Bootstrapped training for tracking-any-point
Carl Doersch, Yi Yang, Dilara Gokay, Pauline Luc, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ross Goroshin, João Carreira, and Andrew Zisserman · 2024
Later among the works it cites.
MemFlow: Optical flow estimation and prediction with memory
Qiaole Dong and Yanwei Fu · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He · 2022
Cited alongside, same era.
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al · 2022
Cited alongside, same era.
Particle video revisited: Tracking through occlusions using point trajectories
Adam W. Harley, Zhaoyuan Fang, and Katerina Fragkiadaki · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Simplerecon: 3d reconstruction without 3d convolutions
Mohamed Sayed, John Gibson, Jamie Watson, Victor Prisacariu, Michael Firman, and Clément Godard · 2022
Cited alongside, same era.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Cited alongside, same era.
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xiaodong Cun, Yong Zhang, Long Quan, and Ying Shan · 2024
Later among the works it cites.
CoTracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2024
Later among the works it cites.
Repurposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler · 2024
Later among the works it cites.
TAPVid-3D: A benchmark for tracking any point in 3d
Skanda Koppula, Ignacio Rocco, Yi Yang, Joe Heyward, João Carreira, Andrew Zisserman, Gabriel Brostow, and Carl Doersch · 2024
Later among the works it cites.
MoSca: Dynamic gaussian fusion from casual videos via 4D motion scaffolds
Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis · 2024
Later among the works it cites.
Grounding image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jérôme Revaud · 2024
Later among the works it cites.
The surprising effectiveness of diffusion models for optical flow and monocular depth estimation
Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J Fleet · 2024
Later among the works it cites.
Learning temporally consistent video depth from video diffusion priors
Jiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Matteo Poggi, and Yiyi Liao · 2024
Later among the works it cites.
3d reconstruction with spatial memory
Hengyi Wang and Lourdes Agapito · 2024
Later among the works it cites.
Any-point trajectory modeling for policy learning
Chuan Wen, Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel · 2024
Later among the works it cites.
SpatialTracker: Tracking any 2D pixels in 3D space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou · 2024
Later among the works it cites.
Flow as the cross-domain manipulation interface
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song · 2024
Later among the works it cites.
V-JEPA 2: Self-supervised video models enable understanding, prediction and planning
Mahmoud Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, and Nicolas Ballas · 2025
Closest in time.
Depth pro: Sharp monocular metric depth in less than a second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun · 2025
Closest in time.
Seurat: From Moving Points to Depth
Seokju Cho, Jiahui Huang, Seungryong Kim, and Joon-Young Lee · 2025
Closest in time.
MegaSaM: Accurate, Fast, and Robust structure and motion from casual dynamic videos
Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holynski, and Noah Snavely · 2025
Closest in time.
Analytisch-Geometrische Entwicklungen
Julius Plücker · 2025
Closest in time.
SpatialTrackerV2: 3D point tracking made easy
Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Iurii Makarov, Bingyi Kang, Xin Zhu, Hujun Bao, Yujun Shen, and Xiaowei Zhou · 2025
Closest in time.
Tapip3d: Tracking any point in persistent 3d geometry
Bowei Zhang, Lei Ke, Adam W Harley, and Katerina Fragkiadaki · 2025
Closest in time.