Fetching the paper…
Reading the bibliography…
We present a novel model for Tracking Any Point (TAP) that effectively tracks any queried point on any physical surface throughout a video sequence.
A computational theory of human stereo vision
David Marr and Tomaso Poggio · 1979
Earlier work this paper cites.
Determining optical flow
Berthold KP Horn and Brian G Schunck · 1981
Earlier work this paper cites.
An iterative image registration technique with an application to stereo vision
Bruce D Lucas and Takeo Kanade · 1981
Earlier work this paper cites.
Finding trajectories of feature points in a monocular image sequence
Ishwar K Sethi and Ramesh Jain · 1987
Earlier work this paper cites.
Object recognition from local scale-invariant features
David G Lowe · 1999
Earlier work this paper cites.
Feature based methods for structure and motion estimation
Philip HS Torr and Andrew Zisserman · 1999
Earlier work this paper cites.
A flexible new technique for camera calibration
Zhengyou Zhang · 2000
Earlier work this paper cites.
Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman · 2003
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
Thomas Brox, Andrés Bruhn, Nils Papenberg, and Joachim Weickert · 2004
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G Lowe · 2004
Earlier work this paper cites.
Object level grouping for video shots
Josef Sivic, Frederik Schaffalitzky, and Andrew Zisserman · 2004
Earlier work this paper cites.
Strike a pose: Tracking people by finding stylized poses
Deva Ramanan, David A Forsyth, and Andrew Zisserman · 2005
Earlier work this paper cites.
Automatic recognition of facial actions in spontaneous expressions
Marian Stewart Bartlett, Gwen Littlewort, Mark G Frank, Claudia Lainscsek, Ian R Fasel, Javier R Movellan, et al · 2006
Earlier work this paper cites.
Surf: Speeded up robust features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool · 2006
Earlier work this paper cites.
Particle video: Long-range motion estimation using point trajectories
Peter Sand and Seth Teller · 2008
Earlier work this paper cites.
Large displacement optical flow
Thomas Brox, Christoph Bregler, and Jitendra Malik · 2009
Earlier work this paper cites.
Sift flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba · 2010
Earlier work this paper cites.
Kinectfusion: Real-time dense surface mapping and tracking
Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon · 2011
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Towards longer long-range motion trajectories
Michael Rubinstein, Ce Liu, and William Freeman · 2012
Earlier work this paper cites.
Towards understanding action recognition
Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J Black · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Supervised descent method and its applications to face alignment
Xuehan Xiong and Fernando De la Torre · 2013
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Yi Sun, Yuheng Chen, Xiaogang Wang, and Xiaoou Tang · 2014
Earlier work this paper cites.
Real-time continuous pose recovery of human hands using convolutional networks
Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin · 2014
Earlier work this paper cites.
Real-time non-rigid reconstruction using an rgb-d camera
Michael Zollhöfer, Matthias Nießner, Shahram Izadi, Christoph Rehmann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew Fitzgibbon, Charles Loop, Christian Theobalt, et al · 2014
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox · 2015
Earlier work this paper cites.
The first facial landmark tracking in-the-wild challenge: Benchmark and results
Jie Shen, Stefanos Zafeiriou, Grigoris G Chrysos, Jean Kossaifi, Georgios Tzimiropoulos, and Maja Pantic · 2015
Cited alongside, same era.
Human pose estimation with iterative error feedback
Joao Carreira, Pulkit Agrawal, Katerina Fragkiadaki, and Jitendra Malik · 2016
Cited alongside, same era.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Cited alongside, same era.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Cited alongside, same era.
Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors
Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krystian Mikolajczyk · 2017
Cited alongside, same era.
Prnet: Self-supervised learning for partial-to-partial registration
Yue Wang and Justin M Solomon · 2019
Later among the works it cites.
Hierarchical discrete distribution decomposition for match density estimation
Zhichao Yin, Trevor Darrell, and Fisher Yu · 2019
Later among the works it cites.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox · 2019
Later among the works it cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Self-supervised learning of interpretable keypoints from unlabelled videos
Tomas Jakab, Ankush Gupta, Hakan Bilen, and Andrea Vedaldi · 2020
Later among the works it cites.
Tsm: Temporal shift module for efficient and scalable video understanding on edge devices
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Cited alongside, same era.
Flownet 2.0: Evolution of optical flow estimation with deep networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox · 2017
Cited alongside, same era.
Desire: Distant future prediction in dynamic scenes with interacting agents
Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker · 2017
Cited alongside, same era.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Cited alongside, same era.
Optical flow estimation using a spatial pyramid network
Anurag Ranjan and Michael J Black · 2017
Cited alongside, same era.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Cited alongside, same era.
Ji Lin, Chuang Gan, Kuan Wang, and Song Han · 2020
Later among the works it cites.
Keypoints into the future: Self-supervised correspondence in model-based reinforcement learning
Lucas Manuelli, Yunzhu Li, Pete Florence, and Russ Tedrake · 2020
Later among the works it cites.
Mantra: Memory augmented networks for multiple trajectory prediction
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, and Alberto Del Bimbo · 2020
Later among the works it cites.
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan · 2021
Later among the works it cites.
On equivariant and invariant learning of object landmark representations
Zezhou Cheng, Jong-Chyi Su, and Subhransu Maji · 2021
Later among the works it cites.
Cotr: Correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi · 2021
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2021
Later among the works it cites.
Monocular 3d reconstruction of interacting hands via collision-aware factorized refinements
Yu Rong, Jingbo Wang, Ziwei Liu, and Chen Change Loy · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Later among the works it cites.
Autoflow: Learning a better training set for optical flow
Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu · 2021
Later among the works it cites.
Rethinking self-supervised correspondence learning: A video frame-level similarity perspective
Jiarui Xu and Xiaolong Wang · 2021
Later among the works it cites.
Tap-vid: A benchmark for tracking any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adrià Recasens, Lucas Smaira, Yusuf Aytar, João Carreira, Andrew Zisserman, and Yi Yang · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
Particle video revisited: Tracking through occlusions using point trajectories
Adam W Harley, Zhaoyuan Fang, and Katerina Fragkiadaki · 2022
Later among the works it cites.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet · 2022
Later among the works it cites.
Coupled iterative refinement for 6d multi-object pose estimation
Lahav Lipson, Zachary Teed, Ankit Goyal, and Jia Deng · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Later among the works it cites.
Zebrapose: Coarse to fine surface encoding for 6dof object pose estimation
Yongzhi Su, Mahdi Saleh, Torben Fetzer, Jason Rambach, Nassir Navab, Benjamin Busam, Didier Stricker, and Federico Tombari · 2022
Later among the works it cites.
Gmflow: Learning optical flow via global matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao · 2022
Later among the works it cites.