Fetching the paper…
Reading the bibliography…
In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the predicted polygons can be later converted into segmentation masks.
Long short-term memory
Jürgen Schmidhuber, Sepp Hochreiter, et al · 1997
Earlier work this paper cites.
Snakes, shapes, and gradient vector flow
Chenyang Xu and J.L. Prince · 1998
Earlier work this paper cites.
LabelMe: a database and web-based tool for image annotation
Bryan C Russell, Antonio Torralba, Kevin P Murphy, and William T Freeman · 2008
Earlier work this paper cites.
The segmented and annotated IAPR TC-12 benchmark
Hugo Jair Escalante, Carlos A. Hernández, Jesus A. Gonzalez, A. López-López, Manuel Montes, Eduardo F. Morales, L. Enrique Sucar, Luis Villaseñor, and Michael Grubinger · 2010
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Fast R-CNN
Ross Girshick · 2015
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Instance-aware semantic segmentation via multi-task network cascades
Jifeng Dai, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
LocNet: Improving localization accuracy for object detection
Spyros Gidaris and Nikos Komodakis · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
Ronghang Hu, Marcus Rohrbach, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Annotating object instances with a Polygon-RNN
Lluis Castrejon, Kaustav Kundu, Raquel Urtasun, and Sanja Fidler · 2017
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille · 2017
Earlier work this paper cites.
The 2017 DAVIS challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Efficient interactive annotation of segmentation datasets with Polygon-RNN++
David Acuna, Huan Ling, Amlan Kar, and Sanja Fidler · 2018
Earlier work this paper cites.
Cascade R-CNN: Delving into high quality object detection
Zhaowei Cai and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Video object segmentation with language referring expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele · 2018
Earlier work this paper cites.
Brute-force facial landmark analysis with a 140,000-way classifier
Mengtian Li, Laszlo Jeni, and Deva Ramanan · 2018
Earlier work this paper cites.
Referring image segmentation via recurrent refinement networks
Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia · 2018
Cited alongside, same era.
Path aggregation network for instance segmentation
Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia · 2018
Cited alongside, same era.
Dynamic multimodal instance segmentation guided by natural language queries
Edgar Margffoy-Tuay, Juan C Pérez, Emilio Botero, and Pablo Arbeláez · 2018
Cited alongside, same era.
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi · 2018
Cited alongside, same era.
Key-word-aware network for referring expression image segmentation
Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu · 2018
Cited alongside, same era.
MAttNet: Modular attention network for referring expression comprehension
Multi-task collaborative network for joint referring expression comprehension and segmentation
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji · 2020
Later among the works it cites.
Deep snake for real-time instance segmentation
Sida Peng, Wen Jiang, Huaijin Pi, Xiuli Li, Hujun Bao, and Xiaowei Zhou · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Later among the works it cites.
Polarmask: Single shot instance segmentation with polar representation
Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu, Ding Liang, Chunhua Shen, and Ping Luo · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Grounding referring expressions in images by variational context
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang · 2018
Cited alongside, same era.
Parallel attention: A unified framework for visual object discovery through dialogs and queries
Bohan Zhuang, Qi Wu, Chunhua Shen, Ian Reid, and Anton Van Den Hengel · 2018
Cited alongside, same era.
YOLACT: Real-time instance segmentation
Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee · 2019
Cited alongside, same era.
See-through-text grouping for referring image segmentation
Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu · 2019
Cited alongside, same era.
Hybrid task cascade for instance segmentation
Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al · 2019
Cited alongside, same era.
Referring expression object segmentation with caption-aware consistency
Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang, Yen-Yu Lin, and Ming-Hsuan Yang · 2019
Cited alongside, same era.
Coatnet: Marrying convolution and attention for all data sizes
Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan · 2021
Later among the works it cites.
Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang · 2021
Later among the works it cites.
Encoder fusion network with co-attention embedding for referring image segmentation
Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu · 2021
Later among the works it cites.
Locate then segment: A strong pipeline for referring image segmentation
Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan · 2021
Later among the works it cites.
MDETR - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Rethinking positional encoding in language pre-training
Guolin Ke, Di He, and Tie-Yan Liu · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Later among the works it cites.
Referring transformer: A one-step approach to multi-task visual grounding
Muchen Li and Leonid Sigal · 2021
Later among the works it cites.
Chen Liang, Yu Wu, Tianfei Zhou, Wenguan Wang, Zongxin Yang, Yunchao Wei, and Yi Yang · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Bottom-up shift and reasoning for referring image segmentation
Sibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou, and Yizhou Yu · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
X-DETR: A Versatile Architecture for Instance-wise Vision-Language Tasks
Zhaowei Cai, Gukyeong Kwon, Avinash Ravichandran, Erhan Bas, Zhuowen Tu, Rahul Bhotika, and Stefano Soatto · 2022
Later among the works it cites.
Pix2seq: A language modeling framework for object detection
Ting Chen, Saurabh Saxena, Lala Li, David J. Fleet, and Geoffrey Hinton · 2022
Later among the works it cites.
A unified sequence interface for vision tasks
Ting Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin, David J. Fleet, and Geoffrey Hinton · 2022
Later among the works it cites.
Restr: Convolution-free referring image segmentation using transformers
Namyup Kim, Dongwon Kim, Cuiling Lan, Wenjun Zeng, and Suha Kwak · 2022
Later among the works it cites.
Instance segmentation with mask-supervised polygonal boundary transformers
Justin Lazarow, Weijian Xu, and Zhuowen Tu · 2022
Later among the works it cites.
Cross-modal progressive comprehension for referring segmentation
Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, and Guanbin Li · 2022
Later among the works it cites.
Unified-IO: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Later among the works it cites.
CRIS: Clip-driven referring image segmentation
Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu · 2022
Later among the works it cites.
SimVLM: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2022
Later among the works it cites.
Language as queries for referring video object segmentation
Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, and Ping Luo · 2022
Later among the works it cites.
Unitab: Unifying text and box outputs for grounded vision-language modeling
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Later among the works it cites.
LAVT: Language-Aware Vision Transformer for Referring Image Segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr · 2022
Later among the works it cites.
SeqTR: A Simple Yet Universal Network for Visual Grounding
Chaoyang Zhu, Yiyi Zhou, Yunhang Shen, Gen Luo, Xingjia Pan, Mingbao Lin, Chao Chen, Liujuan Cao, Xiaoshuai Sun, and Rongrong Ji · 2022
Later among the works it cites.