Fetching the paper…
Reading the bibliography…
Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences.
Capsule-based Object Tracking with Natural Language Specification. In Proceedings of the 29th ACM International Conference on Multimedia . 1948–1956
Ding Ma and Xiangqian Wu. 2021 · 1956
Earlier work this paper cites.
Shili Zhou, Xuhao Jiang, Weimin Tan, Ruian He, and Bo Yan. 2023 · 1974
Earlier work this paper cites.
Determining optical flow
Berthold KP Horn and Brian G Schunck. 1981 · 1981
Earlier work this paper cites.
An iterative image registration technique with an application to stereo vision. In IJCAI’81: 7th international joint conference on Artificial intelligence , Vol. 2. 674–679
Bruce D Lucas and Takeo Kanade. 1981 · 1981
Earlier work this paper cites.
A framework for the robust estimation of optical flow. In 1993 (4th) International Conference on Computer Vision . IEEE, 231–236
Michael J Black and Padmanabhan Anandan. 1993 · 1993
Earlier work this paper cites.
Lucas/Kanade meets Horn/Schunck: Combining local and global optic flow methods
Andrés Bruhn, Joachim Weickert, and Christoph Schnörr. 2005 · 2005
Earlier work this paper cites.
Particle video: Long-range motion estimation using point trajectories
Peter Sand and Seth Teller. 2008 · 2008
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020 · 2010
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks. In Proceedings of the IEEE international conference on computer vision . 2758–2766
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In NIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Accurate optical flow via direct cost volume processing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1289–1297
Jia Xu, René Ranftl, and Vladlen Koltun. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization. In Proceedings of the ICLR
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE conference on computer vision and pattern recognition . 8934–8943
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. 2018 · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Learning to Compose Hypercolumns for Visual Correspondence. In European Conference on Computer Vision
Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. 2020 · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow. In Proceedings of the ECCV . Springer, 402–419
Zachary Teed and Jia Deng. 2020 · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision . 9650–9660
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Cited alongside, same era.
Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5851–5860
Qi Feng, Vitaly Ablavsky, Qinxun Bai, and Stan Sclaroff. 2021 · 2021
Cited alongside, same era.
Prompting for Multi-Modal Tracking. In Proceedings of the 30th ACM International Conference on Multimedia (MM ’22) . Association for Computing Machinery, New York, NY, USA, 3492–3500
Jinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis, and Jingkuan Song. 2022 · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16816–16825
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
Context-TAP: Tracking Any Point Demands Spatial Context Features
Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yitong Dong, Yijin Li, and Hongsheng Li. 2023 · 2023
Later among the works it cites.
TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement
Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ron Mokady, Amir Hertz, and Amit H Bermano. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Cited alongside, same era.
Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13763–13773
Xiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, and Feng Wu. 2021 · 2021
Cited alongside, same era.
TAP-Vid: A Benchmark for Tracking Any Point in a Video. In NeurIPS Datasets Track
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Recasens Continente, Kucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang. 2022 · 2022
Cited alongside, same era.
Divert more attention to vision-language tracking
Mingzhe Guo, Zhipeng Zhang, Heng Fan, and Liping Jing. 2022 · 2022
Cited alongside, same era.
Particle Video Revisited: Tracking Through Occlusions Using Point Trajectories. In ECCV
Adam W Harley, Zhaoyuan Fang, and Katerina Fragkiadaki. 2022 · 2022
Cited alongside, same era.
Flowformer: A transformer architecture for optical flow. In European Conference on Computer Vision . Springer, 668–685
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li. 2022 · 2022
Cited alongside, same era.
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. 2023 · 2023
Later among the works it cites.
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence. In Advances in Neural Information Processing Systems , A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 47500–47510
Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell. 2023a · 2023
Later among the works it cites.
DiffusionTrack: Diffusion Model For Multi-Object Tracking
Run Luo, Zikai Song, Lintao Ma, Jinlin Wei, Wei Yang, and Min Yang. 2023b · 2023
Later among the works it cites.
MFT: Long-Term Tracking of Every Pixel
Michal Neoral, Jonáš Šerỳch, and Jiří Matas. 2023 · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Later among the works it cites.
Compact transformer tracker with correlative masked modeling
Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2023 · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023 · 2023
Later among the works it cites.
Emergent Correspondence from Image Diffusion. In Thirty-seventh Conference on Neural Information Processing Systems
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. 2023 · 2023
Later among the works it cites.
Tracking Everything Everywhere All at Once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. 2023 · 2023
Later among the works it cites.
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment. In Proceedings of the 31st ACM International Conference on Multimedia (MM ’23) . Association for Computing Machinery, New York, NY, USA, 5552–5561
Chunhui Zhang, Xin Sun, Yiqian Yang, Li Liu, Qiong Liu, Xi Zhou, and Yanfeng Wang. 2023 · 2023
Later among the works it cites.
Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 19855–19865
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J Guibas. 2023 · 2023
Later among the works it cites.
Deem: Diffusion models serve as the eyes of large language models for image perception
Run Luo, Yunshui Li, Longze Chen, Wanwei He, Ting-En Lin, Ziqiang Liu, Lei Zhang, Zikai Song, Xiaobo Xia, Tongliang Liu, et al · 2024
Closest in time.