Fetching the paper…
Reading the bibliography…
Referring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects.
H. W. Kuhn, “The hungarian method for the assignment problem,” Nav. Res. Logist. Q. , vol. 2, no. 1-2, pp. 83–97, 1955
1955
Earlier work this paper cites.
R. C. Atkinson and R. M. Shiffrin, “Human memory: A proposed system and its control processes,” in Psychol. Learn. Motiv. , 1968, vol. 2, pp. 89–195
1968
Earlier work this paper cites.
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black, “Towards understanding action recognition,” in ICCV , 2013, pp. 3192–3199
2013
Earlier work this paper cites.
J. Shi, Q. Yan, L. Xu, and J. Jia, “Hierarchical image saliency detection on extended cssd,” TPAMI , vol. 38, no. 4, pp. 717–729, 2015
2015
Earlier work this paper cites.
R. Hu, M. Rohrbach, and T. Darrell, “Segmentation from natural language expressions,” in ECCV , 2016, pp. 108–124
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR , 2016, pp. 11–20
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV , 2016, pp. 69–85
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in CVPR , 2017, pp. 136–145
2017
Earlier work this paper cites.
S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “One-shot video object segmentation,” in CVPR , 2017, pp. 221–230
2017
Earlier work this paper cites.
F. Perazzi, A. Khoreva, R. Benenson, B. Schiele, and A. Sorkine-Hornung, “Learning video object segmentation from static images,” in CVPR , 2017, pp. 2663–2672
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in CVPR , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in ECCV , 2018, pp. 801–818
2018
Earlier work this paper cites.
K.-K. Maninis, S. Caelles, Y. Chen, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “Video object segmentation without temporal information,” TPAMI , vol. 41, no. 6, pp. 1515–1530, 2018
2018
Earlier work this paper cites.
J. Luiten, P. Voigtlaender, and B. Leibe, “Premvos: Proposal-generation, refinement and merging for video object segmentation,” in ACCV , 2018, pp. 565–580
2018
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, K. Sunkavalli, and S. J. Kim, “Fast video object segmentation by reference-guided mask propagation,” in CVPR , 2018, pp. 7376–7385
2018
Earlier work this paper cites.
L. Yang, Y. Wang, X. Xiong, J. Yang, and A. K. Katsaggelos, “Efficient video object segmentation via network modulation,” in CVPR , 2018, pp. 6499–6507
2018
Earlier work this paper cites.
J. Cheng, Y.-H. Tsai, W.-C. Hung, S. Wang, and M.-H. Yang, “Fast and accurate online video object segmentation via tracking parts,” in CVPR , 2018, pp. 7415–7424
2018
Earlier work this paper cites.
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in ECCV , 2018, pp. 3–19
2018
Earlier work this paper cites.
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek, “Actor and action video segmentation from a sentence,” in CVPR , 2018, pp. 5958–5966
2018
Earlier work this paper cites.
A. Khoreva, A. Rohrbach, and B. Schiele, “Video object segmentation with language referring expressions,” in ACCV , 2019, pp. 123–141
2019
Earlier work this paper cites.
Y. Zeng, P. Zhang, J. Zhang, Z. Lin, and H. Lu, “Towards high-resolution salient object detection,” in ICCV , 2019, pp. 7234–7243
2019
Earlier work this paper cites.
Y. Gui, Y. Tian, D.-J. Zeng, Z.-F. Xie, and Y.-Y. Cai, “Reliable and dynamic appearance modeling and label consistency enforcing for fast and coherent video object segmentation with the bilateral grid,” IEEE Trans. Circuits Syst. Video Technol. , vol. 30, no. 12, pp. 4781–4795, 2019
2019
Earlier work this paper cites.
H. Lin, X. Qi, and J. Jia, “Agss-vos: Attention guided single-shot video object segmentation,” in ICCV , 2019, pp. 3949–3957
2019
Earlier work this paper cites.
S. Xu, D. Liu, L. Bao, W. Liu, and P. Zhou, “Mhp-vos: Multiple hypotheses propagation for video object segmentation,” in CVPR , 2019, pp. 314–323
2019
Earlier work this paper cites.
L. Zhang, Z. Lin, J. Zhang, H. Lu, and Y. He, “Fast video object segmentation via dynamic targeting network,” in ICCV , 2019, pp. 5582–5591
2019
Earlier work this paper cites.
P. Voigtlaender, Y. Chai, F. Schroff, H. Adam, B. Leibe, and L.-C. Chen, “Feelvos: Fast end-to-end embedding learning for video object segmentation,” in CVPR , 2019, pp. 9481–9490
2019
Earlier work this paper cites.
Z. Wang, J. Xu, L. Liu, F. Zhu, and L. Shao, “Ranet: Ranking attention network for fast video object segmentation,” in ICCV , 2019, pp. 3978–3987
2019
Earlier work this paper cites.
K. Duarte, Y. S. Rawat, and M. Shah, “Capsulevos: Semi-supervised video object segmentation using capsule routing,” in ICCV , 2019, pp. 8480–8489
2019
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” in ICCV , 2019, pp. 9226–9235
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in CVPR , 2019, pp. 658–666
2019
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in CVPR , 2019, pp. 10 502–10 511
2019
Cited alongside, same era.
H. Wang, C. Deng, J. Yan, and D. Tao, “Asymmetric cross-guided attention network for actor and action video segmentation from natural language query,” in ICCV , 2019, pp. 3939–3948
2019
Cited alongside, same era.
W. Liu, G. Lin, T. Zhang, and Z. Liu, “Guided co-segmentation network for fast video object segmentation,” IEEE Trans. Circuits Syst. Video Technol. , vol. 31, no. 4, pp. 1607–1617, 2020
2020
Cited alongside, same era.
Z. Tan, B. Liu, Q. Chu, H. Zhong, Y. Wu, W. Li, and N. Yu, “Real time video object segmentation in compressed domain,” IEEE Trans. Circuits Syst. Video Technol. , vol. 31, no. 1, pp. 175–188, 2020
2020
Cited alongside, same era.
C. Shang, H. Li, H. Qiu, Q. Wu, F. Meng, T. Zhao, and K. N. Ngan, “Cross-modal recurrent semantic comprehension for referring image segmentation,” IEEE Trans. Circuits Syst. Video Technol. , 2022
2022
Later among the works it cites.
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in ECCV , 2022, pp. 640–658
2022
Later among the works it cites.
D. Li, R. Li, L. Wang, Y. Wang, J. Qi, L. Zhang, T. Liu, Q. Xu, and H. Lu, “You only infer once: Cross-modal meta-transfer for referring video object segmentation,” in AAAI , 2022, pp. 1297–1305
2022
Later among the works it cites.
J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in CVPR , 2022, pp. 4974–4984
2022
Later among the works it cites.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” in CVPR , 2022, pp. 4985–4995
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Z. Yang, Y. Wei, and Y. Yang, “Collaborative video object segmentation by foreground-background integration,” in ECCV , 2020, pp. 332–348
2020
Cited alongside, same era.
H. K. Cheng, J. Chung, Y.-W. Tai, and C.-K. Tang, “Cascadepsp: Toward class-agnostic and very high-resolution segmentation via global and local refinement,” in CVPR , 2020, pp. 8890–8899
2020
Cited alongside, same era.
X. Li, T. Wei, Y. P. Chen, Y.-W. Tai, and C.-K. Tang, “Fss-1000: A 1000-class dataset for few-shot segmentation,” in CVPR , 2020, pp. 2869–2878
2020
Cited alongside, same era.
S. Seo, J.-Y. Lee, and B. Han, “Urvos: Unified referring video object segmentation network with a large-scale benchmark,” in ECCV , 2020, pp. 208–223
2020
Cited alongside, same era.
T. Meinhardt and L. Leal-Taixe, “Make one-shot video object segmentation efficient again,” in NeurIPS , 2020, pp. 10 607–10 619
2020
Cited alongside, same era.
X. Chen, Z. Li, Y. Yuan, G. Yu, J. Shen, and D. Qi, “State-aware tracker for real-time video object segmentation,” in CVPR , 2020, pp. 9384–9393
2020
Cited alongside, same era.
H. Seong, J. Hyun, and E. Kim, “Kernelized memory network for video object segmentation,” in ECCV , 2020, pp. 629–645
2020
Cited alongside, same era.
2022
Later among the works it cites.
X. Xu, J. Zhao, J. Wu, and F. Shen, “Switch and refine: A long-term tracking and segmentation framework,” IEEE Trans. Circuits Syst. Video Technol. , vol. 33, no. 3, pp. 1291–1304, 2022
2022
Later among the works it cites.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Self-supervised video object segmentation by motion-aware mask propagation,” in ICME , 2022, pp. 1–6
2022
Later among the works it cites.
X. Xu, J. Wang, X. Li, and Y. Lu, “Reliable propagation-correction modulation for video object segmentation,” in AAAI , 2022, pp. 2946–2954
2022
Later among the works it cites.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Regional video object segmentation by efficient motion-aware mask propagation,” in DICTA , 2022, pp. 1–6
2022
Later among the works it cites.
Z. Yang and Y. Yang, “Decoupling features in hierarchical propagation for video object segmentation,” NeurIPS , pp. 36 324–36 336, 2022
2022
Later among the works it cites.
W. Zhao, K. Wang, X. Chu, F. Xue, X. Wang, and Y. You, “Modeling motion with multi-modal features for text-based video segmentation,” in CVPR , 2022, pp. 11 737–11 746
2022
Later among the works it cites.
D. Wu, X. Dong, L. Shao, and J. Shen, “Multi-level representation learning with semantic alignment for referring video object segmentation,” in CVPR , 2022, pp. 4996–5005
2022
Later among the works it cites.
Z. Ding, T. Hui, J. Huang, X. Wei, J. Han, and S. Liu, “Language-bridged spatial-temporal interaction for referring video object segmentation,” in CVPR , 2022, pp. 4964–4973
2022
Later among the works it cites.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in CVPR , 2022, pp. 3202–3211
2022
Later among the works it cites.
W. Chen, D. Hong, Y. Qi, Z. Han, S. Wang, L. Qing, Q. Huang, and G. Li, “Multi-attention network for compressed video referring object segmentation,” in ACM MM , 2022, pp. 4416–4425
2022
Later among the works it cites.
H. Ding, C. Liu, S. Wang, and X. Jiang, “Vlt: Vision-language transformer and query generation for referring segmentation,” TPAMI , vol. 45, no. 06, pp. 7900–7916, 2023
2023
Later among the works it cites.
H. Li, M. Sun, J. Xiao, E. G. Lim, and Y. Zhao, “Fully and weakly supervised referring expression segmentation with end-to-end learning,” IEEE Trans. Circuits Syst. Video Technol. , 2023
2023
Later among the works it cites.
M. Feng, H. Hou, L. Zhang, Y. Guo, H. Yu, Y. Wang, and A. Mian, “Exploring hierarchical spatial layout cues for 3d point cloud based scene graph prediction,” IEEE Transactions on Multimedia , 2023
2023
Later among the works it cites.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Spectrum-guided multi-granularity referring video object segmentation,” in ICCV , 2023, pp. 920–930
2023
Later among the works it cites.
Y. Lu, J. Zhang, S. Sun, Q. Guo, Z. Cao, S. Fei, B. Yang, and Y. Chen, “Label-efficient video object segmentation with motion clues,” IEEE Trans. Circuits Syst. Video Technol. , 2023
2023
Later among the works it cites.
Y. Chen, D. Zhang, Y. Zheng, Z.-X. Yang, E. Wu, and H. Zhao, “Boosting video object segmentation via robust and efficient memory network,” IEEE Trans. Circuits Syst. Video Technol. , 2023
2023
Later among the works it cites.
F. Lin, Z. Qiu, C. Liu, T. Yao, H. Xie, and Y. Zhang, “Prototypical matching networks for video object segmentation,” IEEE Transactions on Image Processing , 2023
2023
Later among the works it cites.
J. Tang, G. Zheng, and S. Yang, “Temporal collection and distribution for referring video object segmentation,” in ICCV , 2023, pp. 15 466–15 476
2023
Later among the works it cites.
M. Sun, J. Xiao, E. G. Lim, C. Zhao, and Y. Zhao, “Unified multi-modality video object segmentation using reinforcement learning,” IEEE Trans. Circuits Syst. Video Technol. , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Li, J. Wang, X. Xu, X. Li, B. Raj, and Y. Lu, “Robust referring video object segmentation with cyclic structural consensus,” in ICCV , 2023, pp. 22 236–22 245
2023
Later among the works it cites.
D. Wu, T. Wang, Y. Zhang, X. Zhang, and J. Shen, “Onlinerefer: A simple online baseline for referring video object segmentation,” in ICCV , 2023, pp. 2761–2770
2023
Later among the works it cites.
M. Han, Y. Wang, Z. Li, L. Yao, X. Chang, and Y. Qiao, “Html: Hybrid temporal-scale multimodal learning framework for referring video object segmentation,” in ICCV , 2023, pp. 13 414–13 423
2023
Later among the works it cites.
Z. Luo, Y. Xiao, Y. Liu, S. Li, Y. Wang, Y. Tang, X. Li, and Y. Yang, “Soc: Semantic-assisted object cluster for referring video object segmentation,” NeurIPS , vol. 36, 2023
2023
Later among the works it cites.
M. Gao, J. Yang, J. Han, K. Lu, F. Zheng, and G. Montana, “Decoupling multimodal transformers for referring video object segmentation,” IEEE Trans. Circuits Syst. Video Technol. , 2023
2023
Later among the works it cites.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “Mevis: A large-scale benchmark for video segmentation with motion expressions,” in ICCV , 2023, pp. 2694–2703
2023
Later among the works it cites.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Region aware video object segmentation with deep motion modeling,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
2024
Closest in time.
S. He and H. Ding, “Decoupling static and hierarchical motion perception for referring video segmentation,” in CVPR , 2024
2024
Closest in time.
B. Miao, L. Zhou, A. S. Mian, T. L. Lam, and Y. Xu, “Object-to-scene: Learning to transfer object knowledge to indoor scene recognition,” in IROS . IEEE, 2021, pp. 2069–2075
2075
Closest in time.