Fetching the paper…
Reading the bibliography…
Video object segmentation (VOS) aims to distinguish and track target objects in a video.
I. Cohen and G. Medioni, “Detecting and tracking moving objects for video surveillance,” in Proceedings. 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No PR00149) , vol. 2. IEEE, 1999, pp. 319–325
1999
Earlier work this paper cites.
P. Sand and S. Teller, “Particle video: Long-range motion estimation using point trajectories,” International journal of computer vision , vol. 80, pp. 72–91, 2008
2008
Earlier work this paper cites.
K. N. Ngan and H. Li, Video segmentation and its applications . Springer Science & Business Media, 2011
2011
Earlier work this paper cites.
A. Prest, C. Leistner, J. Civera, C. Schmid, and V. Ferrari, “Learning object class detectors from weakly annotated video,” in 2012 IEEE Conference on computer vision and pattern recognition . IEEE, 2012, pp. 3282–3289
2012
Earlier work this paper cites.
P. Ochs, J. Malik, and T. Brox, “Segmentation of moving objects by long term video analysis,” IEEE transactions on pattern analysis and machine intelligence , vol. 36, no. 6, pp. 1187–1200, 2013
2013
Earlier work this paper cites.
A. Erdélyi, T. Barát, P. Valet, T. Winkler, and B. Rinner, “Adaptive cartooning for privacy protection in camera networks,” in 2014 11th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS) . IEEE, 2014, pp. 44–49
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
G. Ros, S. Ramos, M. Granados, A. Bakhtiary, D. Vazquez, and A. M. Lopez, “Vision-based offline-online perception paradigm for autonomous driving,” in 2015 IEEE Winter Conference on Applications of Computer Vision . IEEE, 2015, pp. 231–238
2015
Earlier work this paper cites.
Z. Zhang, S. Fidler, and R. Urtasun, “Instance-level segmentation for autonomous driving with deep densely connected mrfs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 669–677
2016
Earlier work this paper cites.
K. Saleh, M. Hossny, and S. Nahavandi, “Kangaroo vehicle collision detection using deep semantic segmentation convolutional neural network,” in 2016 International Conference on Digital Image Computing: Techniques and Applications (DICTA) . IEEE, 2016, pp. 1–7
2016
Earlier work this paper cites.
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 724–732
2016
Earlier work this paper cites.
M. Mueller, N. Smith, and B. Ghanem, “A benchmark and simulator for uav tracking,” in European conference on computer vision . Springer, 2016, pp. 445–461
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “One-shot video object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 221–230
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Shin Yoon, F. Rameau, J. Kim, S. Lee, S. Shin, and I. So Kweon, “Pixel-level matching for video object segmentation using convolutional neural networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2167–2176
2017
Earlier work this paper cites.
F. Perazzi, A. Khoreva, R. Benenson, B. Schiele, and A. Sorkine-Hornung, “Learning video object segmentation from static images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2663–2672
2017
Earlier work this paper cites.
W.-D. Jang and C.-S. Kim, “Online video object segmentation via convolutional trident network,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 5849–5858
2017
Earlier work this paper cites.
Y.-T. Hu, J.-B. Huang, and A. Schwing, “Maskrnn: Instance level video object segmentation,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Khoreva, R. Benenson, E. Ilg, T. Brox, and B. Schiele, “Lucid data dreaming for object tracking,” in The DAVIS challenge on video object segmentation , 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, K. Sunkavalli, and S. J. Kim, “Fast video object segmentation by reference-guided mask propagation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7376–7385
2018
Earlier work this paper cites.
N. Xu, L. Yang, Y. Fan, J. Yang, D. Yue, Y. Liang, B. Price, S. Cohen, and T. Huang, “Youtube-vos: Sequence-to-sequence video object segmentation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 585–601
2018
Earlier work this paper cites.
Y. Chen, J. Pont-Tuset, A. Montes, and L. Van Gool, “Blazingly fast video object segmentation with pixel-wise metric learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1189–1198
2018
Earlier work this paper cites.
J. Valmadre, L. Bertinetto, J. F. Henriques, R. Tao, A. Vedaldi, A. W. Smeulders, P. H. Torr, and E. Gavves, “Long-term tracking in the wild: A benchmark,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 670–685
2018
Earlier work this paper cites.
K.-K. Maninis, S. Caelles, Y. Chen, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “Video object segmentation without temporal information,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 6, pp. 1515–1530, 2018
2018
Earlier work this paper cites.
H. Xiao, J. Feng, G. Lin, Y. Liu, and M. Zhang, “Monet: Deep motion exploitation for video object segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1140–1148
2018
Earlier work this paper cites.
Y.-T. Hu, J.-B. Huang, and A. G. Schwing, “Videomatch: Matching based video object segmentation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 54–70
2018
Earlier work this paper cites.
L. Bao, B. Wu, and W. Liu, “Cnn in mrf: Video object segmentation via inference in a cnn-based higher-order spatio-temporal mrf,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5977–5986
2018
Earlier work this paper cites.
J. Cheng, Y.-H. Tsai, W.-C. Hung, S. Wang, and M.-H. Yang, “Fast and accurate online video object segmentation via tracking parts,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7415–7424
2018
Earlier work this paper cites.
P. Hu, G. Wang, X. Kong, J. Kuen, and Y.-P. Tan, “Motion-guided cascaded refinement network for video object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1400–1409
2018
Earlier work this paper cites.
X. Li and C. C. Loy, “Video object segmentation with joint re-identification and attention-aware mask propagation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 90–105
2018
Earlier work this paper cites.
L. Yang, Y. Wang, X. Xiong, J. Yang, and A. K. Katsaggelos, “Efficient video object segmentation via network modulation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6499–6507
2018
Earlier work this paper cites.
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek, “Actor and action video segmentation from a sentence,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5958–5966
2018
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Fast user-guided video object segmentation by interaction-and-propagation networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5247–5256
2019
Earlier work this paper cites.
P. Voigtlaender, Y. Chai, F. Schroff, H. Adam, B. Leibe, and L.-C. Chen, “Feelvos: Fast end-to-end embedding learning for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9481–9490
2019
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9226–9235
2019
Earlier work this paper cites.
L. Yang, Y. Fan, and N. Xu, “Video instance segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5188–5197
2019
Earlier work this paper cites.
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, R. Pflugfelder, J.-K. Kamarainen, L. ˇCehovin Zajc, O. Drbohlav, A. Lukezic, A. Berg et al. , “The seventh visual object tracking vot2019 challenge results,” in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops , 2019, pp. 0–0
2019
Earlier work this paper cites.
H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y. Xu, C. Liao, and H. Ling, “Lasot: A high-quality benchmark for large-scale single object tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 5374–5383
2019
Earlier work this paper cites.
W. Wang, H. Song, S. Zhao, J. Shen, S. Zhao, S. C. Hoi, and H. Ling, “Learning unsupervised video object segmentation through visual attention,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3064–3074
2019
Earlier work this paper cites.
X. Lu, W. Wang, C. Ma, J. Shen, L. Shao, and F. Porikli, “See more, know more: Unsupervised video object segmentation with co-attention siamese networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3623–3632
2019
Earlier work this paper cites.
W. Wang, X. Lu, J. Shen, D. J. Crandall, and L. Shao, “Zero-shot video object segmentation via attentive graph neural networks,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 9236–9245
2019
Cited alongside, same era.
L. Zhang, Z. Lin, J. Zhang, H. Lu, and Y. He, “Fast video object segmentation via dynamic targeting network,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5582–5591
2019
Cited alongside, same era.
K. Xu, L. Wen, G. Li, L. Bo, and Q. Huang, “Spatiotemporal cnn for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 1379–1388
2019
Cited alongside, same era.
C. Ventura, M. Bellver, A. Girbau, A. Salvador, F. Marques, and X. Giro-i Nieto, “Rvos: End-to-end recurrent network for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5277–5286
2019
H. Seong, S. W. Oh, J.-Y. Lee, S. Lee, S. Lee, and E. Kim, “Hierarchical memory matching network for video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 12 889–12 898
2021
Later among the works it cites.
Y. Mao, N. Wang, W. Zhou, and H. Li, “Joint inductive and transductive learning for video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 9670–9679
2021
Later among the works it cites.
L. Ye, M. Rochan, Z. Liu, X. Zhang, and Y. Wang, “Referring segmentation in images and videos with cross-modal self-attention network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 7, pp. 3719–3732, 2021
2021
Later among the works it cites.
Z. Yin, J. Zheng, W. Luo, S. Qian, H. Zhang, and S. Gao, “Learning to recommend frame for interactive video object segmentation in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 15 445–15 454
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr, “Fast online object tracking and segmentation: A unifying approach,” in Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition , 2019, pp. 1328–1338
2019
Cited alongside, same era.
Z. Wang, J. Xu, L. Liu, F. Zhu, and L. Shao, “Ranet: Ranking attention network for fast video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3978–3987
2019
Cited alongside, same era.
J. Johnander, M. Danelljan, E. Brissman, F. S. Khan, and M. Felsberg, “A generative appearance model for end-to-end video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8953–8962
2019
Cited alongside, same era.
H. Wang, C. Deng, J. Yan, and D. Tao, “Asymmetric cross-guided attention network for actor and action video segmentation from natural language query,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3939–3948
2019
Cited alongside, same era.
Y. Ma, D. Yu, T. Wu, and H. Wang, “Paddlepaddle: An open-source deep learning platform from industrial practice,” Frontiers of Data and Domputing , vol. 1, no. 1, pp. 105–115, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Z. Yang, Y. Wei, and Y. Yang, “Collaborative video object segmentation by foreground-background integration,” in European Conference on Computer Vision . Springer, 2020, pp. 332–348
2020
Cited alongside, same era.
H. Seong, J. Hyun, and E. Kim, “Kernelized memory network for video object segmentation,” in European Conference on Computer Vision . Springer, 2020, pp. 629–645
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Heo, Y. J. Koh, and C.-S. Kim, “Guided interactive video object segmentation using reliability-based attention maps,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7322–7330
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Qi, Y. Gao, Y. Hu, X. Wang, X. Liu, X. Bai, S. Belongie, A. Yuille, P. H. Torr, and S. Bai, “Occluded video instance segmentation: A benchmark,” International Journal of Computer Vision , vol. 130, no. 8, pp. 2022–2039, 2022
2022
Later among the works it cites.
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, J.-K. Kämäräinen, H. J. Chang, M. Danelljan, L. Č. Zajc, A. Lukežič et al. , “The tenth visual object tracking vot2022 challenge results,” in European Conference on Computer Vision . Springer, 2022, pp. 431–460
2022
Later among the works it cites.
L. Yang, Y. Fan, and N. Xu, “The 4th large-scale video object segmentation challenge - video instance segmentation track,” Jun. 2022
2022
Later among the works it cites.
——, “The 4th large-scale video object segmentation challenge - video object segmentation track,” Jun. 2022
2022
Later among the works it cites.
M. Li, L. Hu, Z. Xiong, B. Zhang, P. Pan, and D. Liu, “Recurrent dynamic embedding for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1332–1341
2022
Later among the works it cites.
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in European Conference on Computer Vision . Springer, 2022, pp. 640–658
2022
Later among the works it cites.
Z. Yang and Y. Yang, “Decoupling features in hierarchical propagation for video object segmentation,” Advances in Neural Information Processing Systems , vol. 35, pp. 36 324–36 336, 2022
2022
Later among the works it cites.
G. Pei, F. Shen, Y. Yao, G.-S. Xie, Z. Tang, and J. Tang, “Hierarchical feature alignment network for unsupervised video object segmentation,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXIV . Springer, 2022, pp. 596–613
2022
Later among the works it cites.
——, “Adaptive selection of reference frames for video object segmentation,” IEEE Transactions on Image Processing , vol. 31, pp. 1057–1071, 2022
2022
Later among the works it cites.
K. Park, S. Woo, S. W. Oh, I. S. Kweon, and J.-Y. Lee, “Per-clip video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1352–1361
2022
Later among the works it cites.
Z. Lin, T. Yang, M. Li, Z. Wang, C. Yuan, W. Jiang, and W. Liu, “Swem: Towards real-time video object segmentation with sequential weighted expectation-maximization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1362–1372
2022
Later among the works it cites.
P. Guo, W. Zhang, X. Li, and W. Zhang, “Adaptive online mutual learning bi-decoders for video object segmentation,” IEEE Transactions on Image Processing , vol. 31, pp. 7063–7077, 2022
2022
Later among the works it cites.
W. Zhao, K. Wang, X. Chu, F. Xue, X. Wang, and Y. You, “Modeling motion with multi-modal features for text-based video segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 737–11 746
2022
Later among the works it cites.
Z. Ding, T. Hui, J. Huang, X. Wei, J. Han, and S. Liu, “Language-bridged spatial-temporal interaction for referring video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4964–4973
2022
Later among the works it cites.
J. Yang, Y. Huang, K. Niu, L. Huang, Z. Ma, and L. Wang, “Actor and action modular network for text-based video segmentation,” IEEE Transactions on Image Processing , vol. 31, pp. 4474–4489, 2022
2022
Later among the works it cites.
J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4974–4984
2022
Later among the works it cites.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4985–4995
2022
Later among the works it cites.
2022
Later among the works it cites.
P. Sun, J. Cao, Y. Jiang, Z. Yuan, S. Bai, K. Kitani, and P. Luo, “Dancetrack: Multi-object tracking in uniform appearance and diverse motion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 20 993–21 002
2022
Later among the works it cites.
L. Ke, M. Danelljan, X. Li, Y.-W. Tai, C.-K. Tang, and F. Yu, “Mask transfiner for high-quality instance segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4412–4421
2022
Later among the works it cites.
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 . Springer, 2020, pp. 402–419
2022
Later among the works it cites.
H. Ding, C. Liu, S. He, X. Jiang, P. H. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in ICCV , 2023
2023
Later among the works it cites.
A. Athar, J. Luiten, P. Voigtlaender, T. Khurana, A. Dave, B. Leibe, and D. Ramanan, “Burst: A benchmark for unifying object recognition, segmentation and tracking in video,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 1674–1683
2023
Later among the works it cites.
L. Hong, W. Chen, Z. Liu, W. Zhang, P. Guo, Z. Chen, and W. Zhang, “Lvos: A benchmark for long-term video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 480–13 492
2023
Later among the works it cites.
X. Wei, Y. Bai, Y. Zheng, D. Shi, and Y. Gong, “Autoregressive visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 9697–9706
2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4015–4026
2023
Later among the works it cites.
S. Cho, M. Lee, S. Lee, C. Park, D. Kim, and S. Lee, “Treating motion as option to reduce motion dependency in unsupervised video object segmentation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 5140–5149
2023
Later among the works it cites.
L. Hong, W. Zhang, S. Gao, H. Lu, and W. Zhang, “Simulflow: Simultaneously extracting feature and identifying target for unsupervised video object segmentation,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 7481–7490
2023
Later among the works it cites.
H. Zhao, J. Chen, L. Wang, and H. Lu, “Arkittrack: A new diverse dataset for tracking using mobile rgb-d data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 5126–5135
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Yuan, Y. Wang, L. Wang, X. Zhao, H. Lu, Y. Wang, W. Su, and L. Zhang, “Isomer: Isomerous transformer for zero-shot video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 966–976
2023
Later among the works it cites.
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y. Lee, “Tracking anything with decoupled video segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 1316–1326
2023
Later among the works it cites.