Fetching the paper…
Reading the bibliography…
Referring video object segmentation (RVOS) aims to segment video objects with the guidance of natural language reference.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Actor and action video segmentation from a sentence
Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li, and Cees GM Snoek · 2018
Cited alongside, same era.
Video object segmentation with language referring expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele · 2018
Cited alongside, same era.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Cited alongside, same era.
Hybrid task cascade for instance segmentation
Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al · 2019
Cited alongside, same era.
Deep high-resolution representation learning for human pose estimation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang · 2019
Cited alongside, same era.
Asymmetric cross-guided attention network for actor and action video segmentation from natural language query
Conditional convolutions for instance segmentation
Zhi Tian, Chunhua Shen, and Hao Chen · 2020
Later among the works it cites.
Context modulated dynamic networks for actor and action video segmentation with language queries
Hao Wang, Cheng Deng, Fan Ma, and Yi Yang · 2020
Later among the works it cites.
Collaborative video object segmentation by foreground-background integration
Zongxin Yang, Yunchao Wei, and Yi Yang · 2020
Later among the works it cites.
Resnest: Split-attention networks
Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Haibin Lin, Zhi Zhang, Yue Sun, Tong He, Jonas Mueller, R Manmatha, et al · 2020
Later among the works it cites.
https://youtube-vos.org/challenge/2021/
The 3rd large-scale video object segmentation challenge · 2021
Closest in time.
Deberta: Decoding-enhanced bert with disentangled attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Wang, Cheng Deng, Junchi Yan, and Dacheng Tao · 2019
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Polar relative positional encoding for video-language segmentation
Ke Ning, Lingxi Xie, Fei Wu, and Qi Tian · 2020
Cited alongside, same era.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Cited alongside, same era.
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
Clawcranenet: Leveraging object-level relation for text-based video segmentation
Chen Liang, Yu Wu, Yawei Luo, and Yi Yang · 2021
Closest in time.
Video instance segmentation with a propose-reduce paradigm
Huaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu, and Jiaya Jia · 2021
Closest in time.
Local-global context aware transformer for language-guided video segmentation
Chen Liang, Wenguan Wang, Tianfei Zhou, Jiaxu Miao, Yawei Luo, and Yi Yang · 2023
Closest in time.