Fetching the paper…
Reading the bibliography…
Referring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The segmented and annotated iapr tc-12 benchmark
H. J. Escalante, C. A. Hernández, J. A. Gonzalez, A. López-López, M. Montes, E. F. Morales, L. E. Sucar, L. Villasenor, and M. Grubinger · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Interactive text2pickup networks for natural language-based human–robot collaboration
H. Ahn, S. Choi, N. Kim, G. Cha, and S. Oh · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. Van den Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2019
Earlier work this paper cites.
Exploring the limitations of behavior cloning for autonomous driving
F. Codevilla, E. Santana, A. M. López, and A. Gaidon · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Bi-directional relationship inferring network for referring image segmentation
Locate then segment: A strong pipeline for referring image segmentation
Y. Jing, T. Kong, W. Wang, L. Wang, L. Li, and T. Tan · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
W. Kim, B. Son, and I. Kim · 2021
Later among the works it cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2021
Later among the works it cites.
Mail: A unified mask-image-language trimodal network for referring image segmentation
Z. Li, M. Wang, J. Mei, and Y. Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Hu, G. Feng, J. Sun, L. Zhang, and H. Lu · 2020
Cited alongside, same era.
Referring image segmentation via cross-modal progressive comprehension
S. Huang, T. Hui, S. Liu, G. Li, Y. Wei, J. Han, L. Liu, and B. Li · 2020
Cited alongside, same era.
Linguistic structure guided context modeling for referring image segmentation
T. Hui, S. Liu, S. Huang, G. Li, S. Yu, F. Zhang, and J. Han · 2020
Cited alongside, same era.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
G. Li, N. Duan, Y. Fang, M. Gong, and D. Jiang · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Cited alongside, same era.
Cascade grouped attention network for referring expression segmentation
G. Luo, Y. Zhou, R. Ji, X. Sun, J. Su, C.-W. Lin, and Q. Tian · 2020
Cited alongside, same era.
Multi-task collaborative network for joint referring expression comprehension and segmentation
G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji · 2020
Cited alongside, same era.
End-to-end model-free reinforcement learning for urban driving using implicit affordances
M. Toromanoff, E. Wirbel, and F. Moutarde · 2020
Cited alongside, same era.
Editgan: High-precision semantic image editing
H. Ling, K. Kreis, D. Li, S. W. Kim, A. Torralba, and S. Fidler · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Styleclip: Text-driven manipulation of stylegan imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Later among the works it cites.
Natural language for human-robot collaboration: Problems beyond language grounding
S. Pate, W. Xu, Z. Yang, M. Love, S. Ganguri, and L. L. Wong · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu · 2021
Later among the works it cites.
Bottom-up shift and reasoning for referring image segmentation
S. Yang, M. Xia, G. Li, H.-Y. Zhou, and Y. Yu · 2021
Later among the works it cites.
Lavt: Language-aware vision transformer for referring image segmentation
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu · 2021
Later among the works it cites.
Restr: Convolution-free referring image segmentation using transformers
N. Kim, D. Kim, C. Lan, W. Zeng, and S. Kwak · 2022
Closest in time.