Fetching the paper…
Reading the bibliography…
The task of video object segmentation with referring expressions (language-guided VOS) is to, given a linguistic phrase and a video, generate binary masks for the object to which the phrase refers.
Measuring Agreement for Multinomial Data
Mark Davies and Joseph L. Fleiss · 1982
Earlier work this paper cites.
Speaking: From Intention to Articulation
William J. M. Levelt · 1989
Earlier work this paper cites.
A fast algorithm for the generation of referring expressions
Ehud Reiter and Robert Dale · 1992
Earlier work this paper cites.
GermaNet - a lexical-semantic net for German
Birgit Hamp and Helmut Feldweg · 1997
Earlier work this paper cites.
The prevalence of descriptive referring expressions in news and narrative
Raquel Hervás and Mark Finlayson · 2010
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Philipp Krähenbühl and Vladlen Koltun · 2011
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Can humans fly? action understanding with multiple classes of actors
Chenliang Xu, Shao-Hang Hsieh, Caiming Xiong, and Jason J Corso · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
Ronghang Hu, Marcus Rohrbach, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Varun K. Nagaraja, Vlad I. Morariu, and Larry S. Davis · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Cited alongside, same era.
How redundant are redundant color adjectives? an efficiency-based analysis of color overspecification
Paula Rubio-Fernández · 2016
Cited alongside, same era.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Rethinking atrous convolution for semantic image segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam · 2017
Cited alongside, same era.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel · 2018
Later among the works it cites.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Later among the works it cites.
Searching for ambiguous objects in videos using relational referring expressions
Hazan Anayurt, Sezai Artun Ozyegin, Ulfet Cetin, Utku Aktas, and Sinan Kalkan · 2019
Later among the works it cites.
See-through-text grouping for referring image segmentation
Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Tracking by natural language specification
Zhenyang Li, Ran Tao, Efstratios Gavves, Cees GM Snoek, and Arnold WM Smeulders · 2017
Cited alongside, same era.
Recurrent multimodal interaction for referring image segmentation
Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille · 2017
Cited alongside, same era.
Referring expression generation and comprehension via attributes
Jingyu Liu, Liang Wang, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Cited alongside, same era.
A joint speaker-listener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Cited alongside, same era.
Real-time referring expression comprehension by single-stage grounding network
Xinpeng Chen, Lin Ma, Jingyuan Chen, Zequn Jie, Wei Liu, and Jiebo Luo · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Learning to assemble neural module tree networks for visual grounding
Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha · 2019
Later among the works it cites.
Asymmetric cross-guided attention network for actor and action video segmentation from natural language query
Hao Wang, Cheng Deng, Junchi Yan, and Dacheng Tao · 2019
Later among the works it cites.
Dynamic graph attention for referring expression comprehension
Sibei Yang, Guanbin Li, and Yizhou Yu · 2019
Later among the works it cites.
Cross-modal self-attention network for referring image segmentation
Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang · 2019
Later among the works it cites.
Referring Expression Comprehension with Semantic Visual Relationship and Word Mapping
Chao Zhang, Weiming Li, Wanli Ouyang, Qiang Wang, Woo-Shik Kim, and Sunghoon Hong · 2019
Later among the works it cites.
Real-time visual object tracking with natural language description
Qi Feng, Vitaly Ablavsky, Qinxun Bai, Guorong Li, and Stan Sclaroff · 2020
Closest in time.
Bi-directional relationship inferring network for referring image segmentation
Zhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang, and Huchuan Lu · 2020
Closest in time.
Referring image segmentation via cross-modal progressive comprehension
Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, and Bo Li · 2020
Closest in time.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Closest in time.