Fetching the paper…
Reading the bibliography…
As an important step towards visual reasoning, visual grounding (e.g., phrase localization, referring expression comprehension/segmentation) has been widely explored Previous approaches to referring expression comprehension (REC) or segmentation (RES) either suffer from limited performance, due to a two-stage setup, or require the designing of complex task-specific one-stage architectures.
Improving referring expression grounding with cross-modal attention-guided erasing
X. Liu, Z. Wang, J. Shao, X. Wang, and H. Li · 1959
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
F. Milletari, N. Navab, and S.-A. Ahmadi · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Query-guided regression network with context policy for phrase grounding
K. Chen, R. Kovvuri, and R. Nevatia · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2017
Earlier work this paper cites.
Real-time referring expression comprehension by single-stage grounding network
X. Chen, L. Ma, J. Chen, Z. Jie, W. Liu, and J. Luo · 2018
Cited alongside, same era.
Dynamic multimodal instance segmentation guided by natural language queries
E. Margffoy-Tuay, J. C. Pérez, E. Botero, and P. Arbeláez · 2018
Cited alongside, same era.
Conditional image-text embedding networks
B. A. Plummer, P. Kordas, M. H. Kiapour, S. Zheng, R. Piramuthu, and S. Lazebnik · 2018
Cited alongside, same era.
Yolov3: An incremental improvement
J. Redmon and A. Farhadi · 2018
Cited alongside, same era.
Learning two-branch neural networks for image-text matching tasks
L. Wang, Y. Li, J. Huang, and S. Lazebnik · 2018
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation
L. Ye, M. Rochan, Z. Liu, and Y. Wang · 2019
Later among the works it cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Later among the works it cites.
Large-scale adversarial training for vision-and-language representation learning
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu · 2020
Later among the works it cites.
Bi-directional relationship inferring network for referring image segmentation
Z. Hu, G. Feng, J. Sun, L. Zhang, and H. Lu · 2020
Later among the works it cites.
Referring image segmentation via cross-modal progressive comprehension
S. Huang, T. Hui, S. Liu, G. Li, Y. Wei, J. Han, L. Liu, and B. Li · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Bajaj, L. Wang, and L. Sigal · 2019
Cited alongside, same era.
Referring expression object segmentation with caption-aware consistency
Y. W. Chen, Y. H. Tsai, T. Wang, Y. Y. Lin, and M. H. Yang · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Neural sequential phrase grounding (seqground)
P. Dogan, L. Sigal, and M. Gross · 2019
Cited alongside, same era.
Centernet: Keypoint triplets for object detection
K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian · 2019
Cited alongside, same era.
Learning to compose and reason with language tree structures for visual grounding
R. Hong, D. Liu, X. Mo, X. He, and H. Zhang · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Later among the works it cites.
A real-time cross-modality correlation filtering method for referring expression comprehension
Y. Liao, S. Liu, G. Li, F. Wang, Y. Chen, C. Qian, and B. Li · 2020
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Later among the works it cites.
Improving one-stage visual grounding by recursive sub-query construction
Z. Yang, T. Chen, L. Wang, and J. Luo · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
F. Yu, J. Tang, W. Yin, Y. Sun, H. Tian, H. Wu, and H. Wang · 2020
Later among the works it cites.
TransVG: End-to-end visual grounding with transformers
J. Deng, Z. Yang, T. Chen, W. Zhou, and H. Li · 2021
Closest in time.
Visual grounding with transformers
Y. Du, Z. Fu, Q. Liu, and Y. Wang · 2021
Closest in time.
Fast convergence of detr with spatially modulated co-attention
P. Gao, M. Zheng, X. Wang, J. Dai, and H. Li · 2021
Closest in time.
Locate then segment: A strong pipeline for referring image segmentation
Y. Jing, T. Kong, W. Wang, L. Wang, L. Li, and T. Tan · 2021
Closest in time.
MDETR–modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Closest in time.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2021
Closest in time.