Fetching the paper…
Reading the bibliography…
Pixel-level Video Understanding in the Wild Challenge (PVUW) focus on complex video understanding.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Context contrasted feature and gated multi-scale aggregation for scene segmentation
Henghui Ding, Xudong Jiang, Bing Shuai, Ai Qun Liu, and Gang Wang · 2018
Earlier work this paper cites.
Youtube-vos: A large-scale video object segmentation benchmark
Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas S. Huang · 2018
Earlier work this paper cites.
Phraseclick: toward achieving flexible interactive segmentation by phrase and click
Henghui Ding, Scott Cohen, Brian Price, and Xudong Jiang · 2020
Earlier work this paper cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Earlier work this paper cites.
Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Earlier work this paper cites.
Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
Ho Kei Cheng and Alexander G Schwing · 2022
Earlier work this paper cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Earlier work this paper cites.
Putting the object back into video object segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price, Joon-Young Lee, and Alexander Schwing · 2023
Cited alongside, same era.
MeViS: A large-scale benchmark for video segmentation with motion expressions
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy · 2023
Cited alongside, same era.
MOSE: A new dataset for video object segmentation in complex scenes
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Philip HS Torr, and Song Bai · 2023
Cited alongside, same era.
VLT: Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang · 2023
Cited alongside, same era.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Cited alongside, same era.
Bin Cao, Yisi Zhang, Xuanxu Lin, Xingjian He, Bo Zhao, and Jing Liu · 2024
Closest in time.
Learning better video query with sam for video instance segmentation
Hao Fang, Tong Zhang, Xiaofei Zhou, and Xinxin Zhang · 2024
Closest in time.
Mingqi Gao, Jingnan Luo, Jinyu Yang, Jungong Han, and Feng Zheng · 2024
Closest in time.
Decoupling static and hierarchical motion perception for referring video segmentation
Shuting He and Henghui Ding · 2024
Closest in time.
Segment anything in high quality
Lei Ke, Mingqiao Ye, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangtai Li, Henghui Ding, Wenwei Zhang, Haobo Yuan, Jiangmiao Pang, Guangliang Cheng, Kai Chen, Ziwei Liu, and Chen Change Loy · 2023
Cited alongside, same era.
GRES: generalized referring expression segmentation
Chang Liu, Henghui Ding, and Xudong Jiang · 2023
Cited alongside, same era.
Multi-modal mutual attention and iterative interaction for referring image segmentation
Chang Liu, Henghui Ding, Yulun Zhang, and Xudong Jiang · 2023
Cited alongside, same era.
Instance-specific feature propagation for referring segmentation
Chang Liu, Xudong Jiang, and Henghui Ding · 2023
Cited alongside, same era.
Codalab competitions: An open source platform to organize scientific challenges
Adrien Pavao, Isabelle Guyon, Anne-Catherine Letournel, Dinh-Tuan Tran, Xavier Baro, Hugo Jair Escalante, Sergio Escalera, Tyler Thomas, and Zhen Xu · 2023
Cited alongside, same era.
Dvis: Decoupled video instance segmentation framework
Tao Zhang, Xingye Tian, Yu Wu, Shunping Ji, Xuebo Wang, Yuan Zhang, and Pengfei Wan · 2023
Cited alongside, same era.
Xinyu Liu, Jing Zhang, Kexin Zhang, Yuting Yang, Licheng Jiao, and Shuyuan Yang · 2024
Closest in time.
1st place solution for mose track in cvpr 2024 pvuw workshop: Complex video object segmentation
Deshui Miao, Xin Li, Zhenyu He, Yaowei Wang, and Ming-Hsuan Yang · 2024
Closest in time.
Feiyu Pan, Hao Fang, and Xiankai Lu · 2024
Closest in time.
Towards open vocabulary learning: A survey
Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, et al · 2024
Closest in time.
2nd place solution for mose track in cvpr 2024 pvuw workshop: Complex video object segmentation
Zhensong Xu, Jiangtao Yao, Chengjing Wu, Ting Liu, and Luoqi Liu · 2024
Closest in time.
Referred by multi-modality: A unified temporal transformer for video object segmentation
Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen, Wei Zhang, Hongyang Li, Yu Qiao, Hao Dong, Zhongjiang He, and Peng Gao · 2024
Closest in time.