Fetching the paper…
Reading the bibliography…
Tracking pixels in videos is typically studied as an optical flow estimation problem, where every pixel is described with a displacement vector that locates it in the next frame.
Lucas, B.D., Kanade, T., et al.: An iterative image registration technique with an application to stereo vision, vol. 81. Vancouver (1981)
1981
Earlier work this paper cites.
Tomasi, C., Kanade, T.: Detection and tracking of point. Int J Comput Vis 9
1991
Earlier work this paper cites.
Bregler, C., Hertzmann, A., Biermann, H.: Recovering non-rigid 3d shape from image streams. In: Proceedings IEEE Conference on Computer Vision and Pattern Recognition. CVPR 2000 (Cat. No.PR00662). vol. 2, pp. 690–696 vol.2 (2000)
2000
Earlier work this paper cites.
Sidenbladh, H., Black, M.J., Fleet, D.J.: Stochastic tracking of 3d human figures using 2d image motion. In: European conference on computer vision. pp. 702–718. Springer (2000)
2000
Earlier work this paper cites.
Matthews, L., Ishikawa, T., Baker, S.: The template update problem. IEEE transactions on pattern analysis and machine intelligence 26
2004
Earlier work this paper cites.
Zhao, T., Nevatia, R.: Tracking multiple humans in crowded environment. In: Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004. vol. 2, pp. II–II. IEEE (2004)
2004
Earlier work this paper cites.
Sand, P., Teller, S.: Particle video: Long-range motion estimation using point trajectories. In: CVPR. vol. 2, pp. 2195–2202 (2006)
2006
Earlier work this paper cites.
Salgado, A., Sánchez, J.: Temporal constraints in large optical flow estimation. In: International Conference on Computer Aided Systems Theory. pp. 709–716. Springer (2007)
2007
Earlier work this paper cites.
Sundaram, N., Brox, T., Keutzer, K.: Dense point trajectories by GPU-accelerated large displacement optical flow. In: ECCV (2010)
2010
Earlier work this paper cites.
Brox, T., Malik, J.: Large displacement optical flow: Descriptor matching in variational motion estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 33
2011
Earlier work this paper cites.
Geiger, A., Lenz, P., Stiller, C., Urtasun, R.: Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR) (2013)
2013
Earlier work this paper cites.
Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., Brox, T.: Flownet: Learning optical flow with convolutional networks. In: ICCV. pp. 2758–2766 (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Mayer, N., Ilg, E., Häusser, P., Fischer, P., Cremers, D., Dosovitskiy, A., Brox, T.: A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In: CVPR (2016)
2016
Earlier work this paper cites.
Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., Brox, T.: Flownet 2.0: Evolution of optical flow estimation with deep networks. In: CVPR (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Taketomi, T., Uchiyama, H., Ikeda, S.: Visual slam algorithms: A survey from 2010 to 2016. IPSJ Transactions on Computer Vision and Applications 9
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in neural information processing systems. pp. 5998–6008 (2017)
2017
Cited alongside, same era.
Biggs, B., Roddick, T., Fitzgibbon, A., Cipolla, R.: Creatures great and SMAL: Recovering the shape and motion of animals from video. In: Asian Conference on Computer Vision. pp. 3–19. Springer (2018)
2018
Cited alongside, same era.
Teed, Z., Deng, J.: RAFT: Recurrent all-pairs field transforms for optical flow. In: European Conference on Computer Vision. pp. 402–419. Springer (2020)
2020
Later among the works it cites.
Teed, Z., Deng, J.: RAFT: Recurrent all-pairs field transforms for optical flow. https://github.com/princeton-vl/RAFT (2020)
2020
Later among the works it cites.
Wang, Q., Zhou, X., Hariharan, B., Snavely, N.: Learning feature descriptors using camera pose supervision. In: Proc. European Conference on Computer Vision (ECCV) (2020)
2020
Later among the works it cites.
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: ICCV (2021)
2021
Later among the works it cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR (2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Janai, J., G”uney, F., Ranjan, A., Black, M.J., Geiger, A.: Unsupervised learning of multi-frame optical flow with occlusions. In: European Conference on Computer Vision (ECCV). vol. Lecture Notes in Computer Science, vol 11220, pp. 713–731. Springer, Cham (Sep 2018)
2018
Cited alongside, same era.
Sun, D., Yang, X., Liu, M.Y., Kautz, J.: PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume. In: CVPR (2018)
2018
Cited alongside, same era.
Valmadre, J., Bertinetto, L., Henriques, J.F., Tao, R., Vedaldi, A., Smeulders, A.W., Torr, P.H., Gavves, E.: Long-term tracking in the wild: A benchmark. In: ECCV. pp. 670–685 (2018)
2018
Cited alongside, same era.
Xu, N., Yang, L., Fan, Y., Yang, J., Yue, D., Liang, Y., Price, B., Cohen, S., Huang, T.: Youtube-vos: Sequence-to-sequence video object segmentation. In: ECCV. pp. 585–601 (2018)
2018
Cited alongside, same era.
Lai, Z., Xie, W.: Self-supervised learning for video correspondence flow. In: BMVC (2019)
2019
Cited alongside, same era.
Novotny, D., Ravi, N., Graham, B., Neverova, N., Vedaldi, A.: C3dpo: Canonical 3d pose networks for non-rigid structure from motion. In: Proceedings of the IEEE International Conference on Computer Vision (2019)
2019
Cited alongside, same era.
Ren, Z., Gallo, O., Sun, D., Yang, M.H., Sudderth, E.B., Kautz, J.: A fusion approach for multi-frame optical flow estimation. In: Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV) (2019)
2019
Cited alongside, same era.
Smith, L.N., Topin, N.: Super-convergence: Very fast training of neural networks using large learning rates. In: Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications. vol. 11006, p. 1100612. International Society for Optics and Photonics (2019)
2019
Cited alongside, same era.
2021
Later among the works it cites.
Jiang, W., Trulls, E., Hosang, J., Tagliasacchi, A., Yi, K.M.: COTR: Correspondence Transformer for Matching Across Images. In: ICCV (2021)
2021
Later among the works it cites.
Kong, C., Lucey, S.: Deep non-rigid structure from motion with missing data. IEEE Transactions on Pattern Analysis & Machine Intelligence 43
2021
Later among the works it cites.
Sundararaman, R., De Almeida Braga, C., Marchand, E., Pettre, J.: Tracking pedestrian heads in dense crowd. In: CVPR. pp. 3865–3875 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Wiles, O., Ehrhardt, S., Zisserman, A.: Co-attention for conditioned image matching. In: CVPR (2021)
2021
Later among the works it cites.
Xu, J., Wang, X.: Rethinking self-supervised correspondence learning: A video frame-level similarity perspective. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 10075–10085 (October 2021)
2021
Later among the works it cites.
Yang, C., Lamdouar, H., Lu, E., Zisserman, A., Xie, W.: Self-supervised video object segmentation by motion grouping. In: ICCV (2021)
2021
Later among the works it cites.
Yang, G., Sun, D., Jampani, V., Vlasic, D., Cole, F., Chang, H., Ramanan, D., Freeman, W.T., Liu, C.: LASR: Learning articulated shape reconstruction from a monocular video. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 15980–15989 (2021)
2021
Later among the works it cites.
Yang, G., Sun, D., Jampani, V., Vlasic, D., Cole, F., Liu, C., Ramanan, D.: Viser: Video-specific surface embeddings for articulated 3d shape reconstruction. In: NeurIPS (2021)
2021
Later among the works it cites.
Germain, H., Lepetit, V., Bourmaud, G.: Visual correspondence hallucination: Towards geometric reasoning. In: ICLR (2022)
2022
Closest in time.