Fetching the paper…
Reading the bibliography…
This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment.
Kuhn, H.W.: The hungarian method for the assignment problem. In: Naval research logistics quarterly (1955)
1955
Earlier work this paper cites.
Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T.: Referitgame: Referring to objects in photographs of natural scenes. In: EMNLP. pp. 787–798 (2014)
2014
Earlier work this paper cites.
Hu, R., Rohrbach, M., Darrell, T.: Segmentation from natural language expressions. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV. pp. 108–124 (2016)
2016
Earlier work this paper cites.
Nagaraja, V.K., Morariu, V.I., Davis, L.S.: Modeling context between objects for referring expression understanding. In: ECCV. pp. 792–807 (2016)
2016
Earlier work this paper cites.
Shelhamer, E., Long, J., Darrell, T.: Fully convolutional networks for semantic segmentation. TPAMI pp. 640–651 (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. NeurIPS 30
2017
Earlier work this paper cites.
Zhang, Y., Yuan, L., Guo, Y., He, Z., Huang, I.A., Lee, H.: Discriminative bimodal networks for visual localization and detection with natural language queries. In: CVPR. pp. 557–566 (2017)
2017
Earlier work this paper cites.
Zhang, Y., Yuan, L., Guo, Y., He, Z., Huang, I., Lee, H.: Discriminative bimodal networks for visual localization and detection with natural language queries. In: CVPR. pp. 1090–1099 (2017)
2017
Earlier work this paper cites.
Gavrilyuk, K., Ghodrati, A., Li, Z., Snoek, C.G.M.: Actor and action video segmentation from a sentence. In: CVPR. pp. 5958–5966 (2018)
2018
Earlier work this paper cites.
Khoreva, A., Rohrbach, A., Schiele, B.: Video object segmentation with language referring expressions. In: ACCV. pp. 123–141 (2018)
2018
Earlier work this paper cites.
Ding, H., Jiang, X., Shuai, B., Liu, A.Q., Wang, G.: Semantic correlation promoted shape-variant context for segmentation. In: CVPR. pp. 8885–8894 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Wang, H., Deng, C., Yan, J., Tao, D.: Asymmetric cross-guided attention network for actor and action video segmentation from natural language query. In: ICCV. pp. 3938–3947 (2019)
2019
Earlier work this paper cites.
Ye, L., Rochan, M., Liu, Z., Wang, Y.: Cross-modal self-attention network for referring image segmentation. In: CVPR. pp. 10502–10511 (2019)
2019
Earlier work this paper cites.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: ECCV. pp. 213–229 (2020)
2020
Earlier work this paper cites.
Hu, Z., Feng, G., Sun, J., Zhang, L., Lu, H.: Bi-directional relationship inferring network for referring image segmentation. In: CVPR. pp. 4423–4432 (2020)
2020
Cited alongside, same era.
Huang, S., Hui, T., Liu, S., Li, G., Wei, Y., Han, J., Liu, L., Li, B.: Referring image segmentation via cross-modal progressive comprehension. In: CVPR. pp. 10488–10497 (2020)
2020
Cited alongside, same era.
Luo, G., Zhou, Y., Sun, X., Cao, L., Wu, C., Deng, C., Ji, R.: Multi-task collaborative network for joint referring expression comprehension and segmentation. In: CVPR. pp. 10034–10043 (2020)
2020
Cited alongside, same era.
McIntosh, B., Duarte, K., Rawat, Y.S., Shah, M.: Visual-textual capsule routing for text-based video segmentation. In: CVPR. pp. 9942–9951 (2020)
2020
Cited alongside, same era.
Ning, K., Xie, L., Wu, F., Tian, Q.: Polar relative positional encoding for video-language segmentation. In: IJCAI. p. 10 (2020)
Ding, Z., Hui, T., Huang, J., Wei, X., Han, J., Liu, S.: Language-bridged spatial-temporal interaction for referring video object segmentation. In: CVPR. pp. 4964–4973 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Li, X., Zhang, W., Pang, J., Chen, K., Cheng, G., Tong, Y., Loy, C.C.: Video k-net: A simple, strong, and unified baseline for video segmentation. In: CVPR. pp. 18847–18857 (2022)
2022
Later among the works it cites.
Liu, S., Hui, T., Huang, S., Wei, Y., Li, B., Li, G.: Cross-modal progressive comprehension for referring segmentation. TPAMI pp. 4761–4775 (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Seo, S., Lee, J., Han, B.: URVOS: unified referring video object segmentation network with a large-scale benchmark. In: ECCV. pp. 208–223 (2020)
2020
Cited alongside, same era.
Wang, H., Deng, C., Ma, F., Yang, Y.: Context modulated dynamic networks for actor and action video segmentation with language queries. In: AAAI. pp. 12152–12159 (2020)
2020
Cited alongside, same era.
Ding, H., Liu, C., Wang, S., Jiang, X.: Vision-language transformer and query generation for referring segmentation. In: ICCV. pp. 16301–16310 (2021)
2021
Cited alongside, same era.
Duke, B., Ahmed, A., Wolf, C., Aarabi, P., Taylor, G.W.: Sstvos: Sparse spatiotemporal transformers for video object segmentation. In: CVPR. pp. 5912–5921 (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Miao, J., Wei, Y., Wu, Y., Liang, C., Li, G., Yang, Y.: Vspw: A large-scale dataset for video scene parsing in the wild. In: CVPR. pp. 4133–4143 (2021)
2021
Cited alongside, same era.
2022
Later among the works it cites.
Liu, Y., Yu, R., Yin, F., Zhao, X., Zhao, W., Xia, W., Yang, Y.: Learning quality-aware dynamic memory for video object segmentation. In: ECCV. pp. 468–486 (2022)
2022
Later among the works it cites.
Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer. In: CVPR. pp. 3192–3201 (2022)
2022
Later among the works it cites.
Miao, J., Wang, X., Wu, Y., Li, W., Zhang, X., Wei, Y., Yang, Y.: Large-scale video panoptic segmentation in the wild: A benchmark. In: CVPR. pp. 21033–21043 (2022)
2022
Later among the works it cites.
Wang, Z., Lu, Y., Li, Q., Tao, X., Guo, Y., Gong, M., Liu, T.: Cris: Clip-driven referring image segmentation. In: CVPR. pp. 11686–11695 (2022)
2022
Later among the works it cites.
Wu, J., Jiang, Y., Sun, P., Yuan, Z., Luo, P.: Language as queries for referring video object segmentation. In: CVPR. pp. 4964–4974 (2022)
2022
Later among the works it cites.
Xu, M., Zhang, Z., Wei, F., Lin, Y., Cao, Y., Hu, H., Bai, X.: A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model. In: ECCV. pp. 736–753 (2022)
2022
Later among the works it cites.
Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., Torr, P.H.S.: Lavt: Language-aware vision transformer for referring image segmentation. In: CVPR. pp. 18134–18144 (2022)
2022
Later among the works it cites.
Ye, L., Rochan, M., Liu, Z., Zhang, X., Wang, Y.: Referring segmentation in images and videos with cross-modal self-attention network. TPAMI pp. 3719–3732 (2022)
2022
Later among the works it cites.
2023
Closest in time.