Fetching the paper…
Reading the bibliography…
In this paper, we propose a simple and strong framework for Tracking Any Point with TRansformers (TAPTR).
Horn, B.K., Schunck, B.G.: Determining Optical Flow. Artificial Intelligence p. 185–203 (Aug 1981)
1981
Earlier work this paper cites.
Black, M., Anandan, P.: A Framework for the Robust Estimation of Optical Flow. In: 1993 (4th) International Conference on Computer Vision (Dec 2002)
2002
Earlier work this paper cites.
Bruhn, A., Weickert, J., Schnörr, C.: Lucas/kanade meets horn/schunck: combining local and global optic flow methods. International Journal of Computer Vision,International Journal of Computer Vision (Feb 2005)
2005
Earlier work this paper cites.
Klinker, F.: Exponential moving average versus moving exponential average. Mathematische Semesterberichte 58
2011
Earlier work this paper cites.
Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Smagt, P.v.d., Cremers, D., Brox, T.: FlowNet: Learning Optical Flow with Convolutional Networks. In: 2015 IEEE International Conference on Computer Vision (ICCV) (Dec 2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., Brox, T.: FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (Jul 2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: NIPS (2017)
2017
Earlier work this paper cites.
Xu, J., Ranftl, R., Koltun, V.: Accurate Optical Flow via Direct Cost Volume Processing. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (Jul 2017)
2017
Earlier work this paper cites.
Guler, R.A., Neverova, N., Kokkinos, I.: DensePose: Dense Human Pose Estimation in the Wild. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (Jun 2018)
2018
Earlier work this paper cites.
Sun, D., Yang, X., Liu, M.Y., Kautz, J.: PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (Jun 2018)
2018
Earlier work this paper cites.
Liang, Z., Guo, Y., Feng, Y., Chen, W., Qiao, L., Zhou, L., Zhang, J., Liu, H.: Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy. IEEE transactions on pattern analysis and machine intelligence 43
2019
Earlier work this paper cites.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-End Object Detection with Transformers. In: European conference on computer vision. pp. 213–229. Springer (2020)
2020
Earlier work this paper cites.
Jabri, A., Owens, A., Efros, A.: Space-time correspondence as a contrastive random walk. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
Lai, Z., Lu, E., Xie, W.: Mast: A memory-augmented self-supervised tracker. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6479–6488 (2020)
2020
Earlier work this paper cites.
Ning, G., Pei, J., Huang, H.: LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. pp. 1034–1035 (2020)
2020
Earlier work this paper cites.
Teed, Z., Deng, J.: RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. pp. 402–419. Springer (2020)
2020
Earlier work this paper cites.
Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: Proc. ECCV (2020)
2020
Earlier work this paper cites.
Wang, J., Zhong, Y., Dai, Y., Zhang, K., Ji, P., Li, H.: Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation. Cornell University - arXiv,Cornell University - arXiv (Oct 2020)
2020
Earlier work this paper cites.
Wang, M., Tighe, J., Modolo, D.: Combining detection and tracking for human pose estimation in videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11088–11096 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Poc. ICCV (2021)
2021
Earlier work this paper cites.
Jiang, S., Lu, Y., Li, H., Hartley, R.: Learning Optical Flow from a Few Matches. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Jun 2021)
2021
Cited alongside, same era.
Jiang, W., Trulls, E., Hosang, J., Tagliasacchi, A., Yi, K.M.: COTR: Correspondence Transformer for Matching Across Images. In: ICCV (2021)
2021
Cited alongside, same era.
Meng, D., Chen, X., Fan, Z., Zeng, G., Li, H., Yuan, Y., Sun, L., Wang, J.: Conditional DETR for Fast Training Convergence. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3651–3660 (2021)
2021
Cited alongside, same era.
Shen, Z., Dai, Y., Rao, Z.: CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 13906–13915 (2021)
2021
Cited alongside, same era.
2023
Later among the works it cites.
Li, F., Jiang, Q., Zhang, H., Ren, T., Liu, S., Zou, X., Xu, H., Li, H., Li, C., Yang, J., Zhang, L., Gao, J.: Visual In-Context Prompting (2023)
2023
Later among the works it cites.
Li, F., Zeng, A., Liu, S., Zhang, H., Li, H., Zhang, L., Ni, L.M.: Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR (2023)
2023
Later among the works it cites.
Li, H., Zhang, H., Zeng, Z., Liu, S., Li, F., Ren, T., Zhang, L.: DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6684–6693 (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xu, H., Yang, J., Cai, J., Zhang, J., Tong, X.: High-Resolution Optical Flow from 1D Attention and Correlation. Cornell University - arXiv,Cornell University - arXiv (Apr 2021)
2021
Cited alongside, same era.
Xu, J., Wang, X.: Rethinking self-supervised correspondence learning: A video frame-level similarity perspective. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 10075–10085 (2021)
2021
Cited alongside, same era.
Zhang, F., Woodford, O.J., Prisacariu, V., Torr, P.H.S.: Separable Flow: Learning Motion Cost Volumes for Optical Flow Estimation. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (Oct 2021)
2021
Cited alongside, same era.
Doersch, C., Gupta, A., Markeeva, L., Recasens, A., Smaira, L., Aytar, Y., Carreira, J., Zisserman, A., Yang, Y.: TAP-Vid: A Benchmark for Tracking Any Point in a Video. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Doersch, C., Gupta, A., Markeeva, L., Recasens, A., Smaira, L., Aytar, Y., Carreira, J., Zisserman, A., Yang, Y.: Tap-vid: A benchmark for tracking any point in a video. arXiv (2022)
2022
Cited alongside, same era.
Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D.J., Gnanapragasam, D., Golemo, F., Herrmann, C., et al.: Kubric: A scalable dataset generator. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3749–3761 (2022)
2022
Cited alongside, same era.
Harley, A.W., Fang, Z., Fragkiadaki, K.: Particle Video Revisited: Tracking Through Occlusions Using Point Trajectories. In: European Conference on Computer Vision. pp. 59–75. Springer (2022)
2022
Cited alongside, same era.
Harley, A.W., Fang, Z., Fragkiadaki, K.: Particle video revisited: Tracking through occlusions using point trajectories. In: Proc. ECCV (2022)
2022
Cited alongside, same era.
Liu, S., Ren, T., Chen, J., Zeng, Z., Zhang, H., Li, F., Li, H., Huang, J., Su, H., Zhu, J., Zhang, L.: Detection Transformer with Stable Matching (2023)
2023
Later among the works it cites.
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., Zhang, L.: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Pan, L., Wang, J., Huang, B., Zhang, J., Wang, H., Tang, X., Wang, Y.: Synthesizing Physically Plausible Human Motions in 3D Scenes (2023)
2023
Later among the works it cites.
Ren, T., Liu, S., Li, F., Zhang, H., Zeng, A., Yang, J., Liao, X., Jia, D., Li, H., Cao, H., Wang, J., Zeng, Z., Qi, X., Yuan, Y., Yang, J., Zhang, L.: detrex: Benchmarking Detection Transformers (2023)
2023
Later among the works it cites.
Ren, T., Yang, J., Liu, S., Zeng, A., Li, F., Zhang, H., Li, H., Zeng, Z., Zhang, L.: A Strong and Reproducible Object Detector with Only Public Datasets (2023)
2023
Later among the works it cites.
Shi, X., Huang, Z., Li, D., Zhang, M., Cheung, K.C., See, S., Qin, H., Dai, J., Li, H.: FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1599–1610 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Vendrow, E., Le, D.T., Cai, J., Rezatofighi, H.: JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4811–4820 (2023)
2023
Later among the works it cites.
Wang, Q., Chang, Y.Y., Cai, R., Li, Z., Hariharan, B., Holynski, A., Snavely, N.: Tracking everything everywhere all at once (2023)
2023
Later among the works it cites.
Xu, G., Wang, X., Ding, X., Yang, X.: Iterative Geometry Encoding Volume for Stereo Matching. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 21919–21928 (2023)
2023
Later among the works it cites.
Yuan, Y., Song, J., Iqbal, U., Vahdat, A., Kautz, J.: PhysDiff: Physics-Guided Human Motion Diffusion Model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 16010–16021 (2023)
2023
Later among the works it cites.
Zhao, H., Zhou, H., Zhang, Y., Chen, J., Yang, Y., Zhao, Y.: High-Frequency Stereo Matching Network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1327–1336 (2023)
2023
Later among the works it cites.
Zheng, Y., Harley, A.W., Shen, B., Wetzstein, G., Guibas, L.J.: Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 19855–19865 (2023)
2023
Later among the works it cites.
2024
Closest in time.
Liu, Z., Li, Y., Okutomi, M.: Global Occlusion-Aware Transformer for Robust Stereo Matching. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 3535–3544 (2024)
2024
Closest in time.
Neoral, M., Šerỳch, J., Matas, J.: MFT: Long-Term Tracking of Every Pixel. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 6837–6847 (2024)
2024
Closest in time.
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., Zeng, Z., Zhang, H., Li, F., Yang, J., Li, H., Jiang, Q., Zhang, L.: Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks (2024)
2024
Closest in time.