Fetching the paper…
Reading the bibliography…
Referring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Video object segmentation with language referring expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A Conneau · 2019
Earlier work this paper cites.
Video object segmentation using space-time memory networks
Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim · 2019
Earlier work this paper cites.
Asymmetric cross-guided attention network for actor and action video segmentation from natural language query
Hao Wang, Cheng Deng, Junchi Yan, and Dacheng Tao · 2019
Earlier work this paper cites.
Refvos: A closer look at referring expressions for video object segmentation
Miriam Bellver, Carles Ventura, Carina Silberer, Ioannis Kazakos, Jordi Torres, and Xavier Giro-i Nieto · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
Visual-textual capsule routing for text-based video segmentation
Bruce McIntosh, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah · 2020
Earlier work this paper cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Earlier work this paper cites.
Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang · 2021
Earlier work this paper cites.
Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
Ho Kei Cheng and Alexander G Schwing · 2022
Earlier work this paper cites.
Language-bridged spatial-temporal interaction for referring video object segmentation
Zihan Ding, Tianrui Hui, Junshi Huang, Xiaoming Wei, Jizhong Han, and Si Liu · 2022
Earlier work this paper cites.
You only infer once: Cross-modal meta-transfer for referring video object segmentation
Dezhuang Li, Ruoqi Li, Lijun Wang, Yifan Wang, Jinqing Qi, Lu Zhang, Ting Liu, Qingquan Xu, and Huchuan Lu · 2022
Earlier work this paper cites.
Xmem++: Production-level video segmentation from few annotated frames
Maksym Bekuzarov, Ariana Bermudez, Joon-Young Lee, and Hao Li · 2023
Cited alongside, same era.
Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang · 2023
Cited alongside, same era.
Mevis: A large-scale benchmark for video segmentation with motion expressions
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy · 2023
Cited alongside, same era.
Html: Hybrid temporal-scale multimodal learning framework for referring video object segmentation
Mingfei Han, Yali Wang, Zhihui Li, Lina Yao, Xiaojun Chang, and Yu Qiao · 2023
Cited alongside, same era.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2024
Later among the works it cites.
Refsam: Efficiently adapting segmenting anything model for referring video object segmentation
Yonglin Li, Jing Zhang, Xiao Teng, Long Lan, and Xinwang Liu · 2024
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Later among the works it cites.
Soc: Semantic-assisted object cluster for referring video object segmentation
Zhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li, Yitong Wang, Yansong Tang, Xiu Li, and Yujiu Yang · 2024
Later among the works it cites.
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S Khan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn better for video object segmentation
Meng Lan, Jing Zhang, Lefei Zhang, and Dacheng Tao · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al · 2023
Cited alongside, same era.
Spectrum-guided multi-granularity referring video object segmentation
Bo Miao, Mohammed Bennamoun, Yongsheng Gao, and Ajmal Mian · 2023
Cited alongside, same era.
Temporal collection and distribution for referring video object segmentation
Jiajin Tang, Ge Zheng, and Sibei Yang · 2023
Cited alongside, same era.
Image as a foreign language: Beit pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2023
Cited alongside, same era.
Onlinerefer: A simple online baseline for referring video object segmentation
Dongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang, and Jianbing Shen · 2023
Cited alongside, same era.
u-llava: Unifying multi-modal tasks via large language model
Jinjin Xu, Liwu Xu, Yuzhe Yang, Xiang Li, Fanyi Wang, Yanchun Xie, Yi-Jie Huang, and Yaqian Li · 2023
Cited alongside, same era.
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al · 2024
Later among the works it cites.
Samrs: Scaling-up remote sensing segmentation dataset with segment anything model
Di Wang, Jing Zhang, Bo Du, Minqiang Xu, Lin Liu, Dacheng Tao, and Liangpei Zhang · 2024
Later among the works it cites.
Hyperseg: Towards universal visual segmentation with large language model
Cong Wei, Yujie Zhong, Haoxian Tan, Yong Liu, Zheng Zhao, Jie Hu, and Yujiu Yang · 2024
Later among the works it cites.
Efficientsam: Leveraged masked image pretraining for efficient segment anything
Yunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang, Fanyi Xiao, Chenchen Zhu, Xiaoliang Dai, Dilin Wang, Fei Sun, Forrest Iandola, et al · 2024
Later among the works it cites.
A simple baseline with single-encoder for referring image segmentation
Seonghoon Yu, Ilchae Jung, Byeongju Han, Taeoh Kim, Yunho Kim, Dongyoon Wee, and Jeany Son · 2024
Later among the works it cites.
Losh: Long-short text joint prediction network for referring video object segmentation
Linfeng Yuan, Miaojing Shi, Zijie Yue, and Qijun Chen · 2024
Later among the works it cites.
Surgicalsam: Efficient class promptable surgical instrument segmentation
Wenxi Yue, Jing Zhang, Kun Hu, Yong Xia, Jiebo Luo, and Zhiyong Wang · 2024
Later among the works it cites.
Evf-sam: Early vision-language fusion for text-prompted segment anything model
Yuxuan Zhang, Tianheng Cheng, Rui Hu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang, et al · 2024
Later among the works it cites.
Exploring pre-trained text-to-video diffusion models for referring video object segmentation
Zixin Zhu, Xuelu Feng, Dongdong Chen, Junsong Yuan, Chunming Qiao, and Gang Hua · 2024
Later among the works it cites.
Semantic and sequential alignment for referring video object segmentation
Feiyu Pan, Hao Fang, Fangkai Li, Yanyu Xu, Yawei Li, Luca Benini, and Xiankai Lu · 2025
Closest in time.
Customized sam 2 for referring remote sensing image segmentation
Fu Rong, Meng Lan, Qian Zhang, and Lefei Zhang · 2025
Closest in time.
Logiczsl: Exploring logic-induced representation for compositional zero-shot learning
Peng Wu, Xiankai Lu, Hao Hu, Yongqin Xian, Jianbing Shen, and Wenguan Wang · 2025
Closest in time.