Fetching the paper…
Reading the bibliography…
We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D).
Three-dimensional scene flow
Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade · 1999
Earlier work this paper cites.
Recovering non-rigid 3d shape from image streams
Christoph Bregler, Aaron Hertzmann, and Henning Biermann · 2000
Earlier work this paper cites.
Nonrigid structure from motion in trajectory space
Ijaz Akhter, Yaser Sheikh, Sohaib Khan, and Takeo Kanade · 2008
Earlier work this paper cites.
3d human motion analysis in monocular video: techniques and challenges
Cristian Sminchisescu · 2008
Earlier work this paper cites.
Clustered pose and nonlinear appearance models for human pose estimation
Sam Johnson and Mark Everingham · 2010
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black · 2013
Earlier work this paper cites.
Parsing ikea objects: Fine pose estimation
Joseph J Lim, Hamed Pirsiavash, and Antonio Torralba · 2013
Earlier work this paper cites.
Model-based pose estimation for rigid objects
Manolis Lourakis and Xenophon Zabulis · 2013
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis · 2013
Earlier work this paper cites.
A survey of optical flow techniques for robotics navigation applications
Haiyang Chao, Yu Gu, and Marcello Napolitano · 2014
Earlier work this paper cites.
Real-time continuous pose recovery of human hands using convolutional networks
Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin · 2014
Earlier work this paper cites.
Object scene flow for autonomous vehicles
Moritz Menze and Andreas Geiger · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Towards viewpoint invariant 3d human pose estimation
Albert Haque, Boya Peng, Zelun Luo, Alexandre Alahi, Serena Yeung, and Li Fei-Fei · 2016
Earlier work this paper cites.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Comparative evaluation of hand-crafted and learned local features
Johannes L Schonberger, Hans Hardmeier, Torsten Sattler, and Marc Pollefeys · 2017
Earlier work this paper cites.
Megadepth: Learning singleview depth prediction from internet photos. ieee
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll · 2018
Earlier work this paper cites.
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan · 2018
Earlier work this paper cites.
Creatures great and SMAL: Recovering the shape and motion of animals from video
Benjamin Biggs, Thomas Roddick, Andrew Fitzgibbon, and Roberto Cipolla · 2019
Cited alongside, same era.
Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos
Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova · 2019
Cited alongside, same era.
Panoptic studio: A massively multiview system for social interaction capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, and Timothy Godisart · 2019
Cited alongside, same era.
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
Jiefeng Li, Can Wang, Hao Zhu, Yihuan Mao, Hao-Shu Fang, and Cewu Lu · 2019
Cited alongside, same era.
Deep rigid instance scene flow
Wei-Chiu Ma, Shenlong Wang, Rui Hu, Yuwen Xiong, and Raquel Urtasun · 2019
Cited alongside, same era.
Improved active speaker detection based on optical flow
Humans in 4d: Reconstructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik · 2023
Later among the works it cites.
Rigidity-aware detection for 6d object pose estimation
Yang Hai, Rui Song, Jiaojiao Li, Mathieu Salzmann, and Yinlin Hu · 2023
Later among the works it cites.
CoTracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2023
Later among the works it cites.
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis · 2023
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chong Huang and Kazuhito Koishida · 2020
Cited alongside, same era.
Consistent video depth estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf · 2020
Cited alongside, same era.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2020
Cited alongside, same era.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al · 2020
Cited alongside, same era.
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang · 2021
Cited alongside, same era.
Neural scene flow fields for space-time view synthesis of dynamic scenes
Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang · 2021
Cited alongside, same era.
Pretraining boosts out-of-domain robustness for pose estimation
Alexander Mathis, Thomas Biasi, Steffen Schneider, Mert Yuksekgonul, Byron Rogers, Matthias Bethge, and Mackenzie W Mathis · 2021
Cited alongside, same era.
Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, and Noah Snavely · 2023
Later among the works it cites.
Centertube: Tracking multiple 3d objects with 4d tubelets in dynamic point clouds
Hao Liu, Yanni Ma, Qingyong Hu, and Yulan Guo · 2023
Later among the works it cites.
Introducing Project Aria Glasses, from Meta, 2023
Meta · 2023
Later among the works it cites.
Aria digital twin: A new benchmark dataset for egocentric 3d machine perception
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Yuheng Carl Ren · 2023
Later among the works it cites.
Any-point trajectory modeling for policy learning
Chuan Wen, Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel · 2023
Later among the works it cites.
Argoverse 2: Next generation datasets for self-driving perception and forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, et al · 2023
Later among the works it cites.
Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting
Zeyu Yang, Hongye Yang, Zijie Pan, Xiatian Zhu, and Li Zhang · 2023
Later among the works it cites.
VideoDoodles: Hand-drawn animations on videos with scene-aware canvases
Emilie Yu, Kevin Blackburn-Matzen, Cuong Nguyen, Oliver Wang, Rubaiat Habib Kazi, and Adrien Bousseau · 2023
Later among the works it cites.
PointOdyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J Guibas · 2023
Later among the works it cites.
Bootstap: Bootstrapped training for tracking-any-point
Carl Doersch, Yi Yang, Dilara Gokay, Pauline Luc, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ross Goroshin, João Carreira, and Andrew Zisserman · 2024
Closest in time.
I can’t believe it’s not scene flow!
Ishan Khatri, Kyle Vedder, Neehar Peri, Deva Ramanan, and James Hays · 2024
Closest in time.
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al · 2024
Closest in time.
RoboTAP: Tracking arbitrary points for few-shot visual imitation
Mel Vecerik, Carl Doersch, Yi Yang, Todor Davchev, Yusuf Aytar, Guangyao Zhou, Raia Hadsell, Lourdes Agapito, and Jon Scholz · 2024
Closest in time.
SceneTracker: Long-term scene flow estimation network
Bo Wang, Jian Li, Yang Yu, Li Liu, Zhenping Sun, and Dewen Hu · 2024
Closest in time.
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Closest in time.