Fetching the paper…
Reading the bibliography…
As an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning.
X. Liu, Z. Wang, J. Shao, X. Wang, and H. Li, “Improving referring expression grounding with cross-modal attention-guided erasing,” in Proc. CVPR , 2019, pp. 1950–1959
1959
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proc. ICML , 2009
2009
Earlier work this paper cites.
M. Kumar, B. Packer, and D. Koller, “Self-paced learning for latent variable models,” in Proc. NeurIPS , 2010
2010
Earlier work this paper cites.
H. J. Escalante, C. A. Hernández, J. A. Gonzalez, A. López-López, M. Montes, E. F. Morales, L. E. Sucar, L. Villasenor, and M. Grubinger, “The segmented and annotated iapr tc-12 benchmark,” Computer vision and image understanding , vol. 114, no. 4, pp. 419–428, 2010
2010
Earlier work this paper cites.
X. Li, A. Dick, H. Wang, C. Shen, and A. van den Hengel, “Graph mode-based contextual kernels for robust svm tracking,” in Proc. ICCV , 2011, pp. 1156–1163
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. NeurIPS , 2012
2012
Earlier work this paper cites.
Y. Chen, X. Li, A. Dick, and R. Hill, “Ranking consistency for image matching and object retrieval,” Pattern Recognition , vol. 47, no. 3, pp. 1349–1360, 2014
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proc. EMNLP , 2014, pp. 787–798
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proc. ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proc. ICCV , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Proc. NeurIPS , 2015
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proc. ICCV , 2015, pp. 2641–2649
2015
Earlier work this paper cites.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in Proc. CVPR , 2016, pp. 4995–5004
2016
Earlier work this paper cites.
A. Salvador, X. Giró-i Nieto, F. Marqués, and S. Satoh, “Faster r-cnn features for instance search,” in Proc. CVPR , 2016, pp. 9–16
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in Proc. ECCV , 2016, pp. 21–37
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Proc. ECCV , 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proc. CVPR , 2016, pp. 11–20
2016
Earlier work this paper cites.
V. K. Nagaraja, V. I. Morariu, and L. S. Davis, “Modeling context between objects for referring expression understanding,” in Proc. ECCV , 2016, pp. 792–807
2016
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in Proc. ECCV , 2016, pp. 817–834
2016
Earlier work this paper cites.
W. Yu, K. Yang, H. Yao, X. Sun, and P. Xu, “Exploiting the complementary strengths of multi-layer cnn features for image retrieval,” Neurocomputing , vol. 237, pp. 235–241, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” Int. J. Comput. Vis. , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proc. CVPR , 2018, pp. 1307–1315
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
Z. Yang, B. Gong, L. Wang, W. Huang, D. Yu, and J. Luo, “A fast and accurate one-stage approach to visual grounding,” in Proc. ICCV , 2019, pp. 4683–4693
2019
Cited alongside, same era.
X. Rong, C. Yi, and Y. Tian, “Unambiguous scene text segmentation with referring expression comprehension,” IEEE Trans. Image Process. , vol. 29, pp. 591–601, 2019
2019
Cited alongside, same era.
X. Jiang, L. Zhang, P. Lv, Y. Guo, R. Zhu, Y. Li, Y. Pang, X. Li, B. Zhou, and M. Xu, “Learning multi-level density maps for crowd counting,” IEEE Trans. Neural Networks and Learning Systems. , vol. 31, no. 8, pp. 2705–2715, 2019
2019
Cited alongside, same era.
M. Li and L. Sigal, “Referring transformer: A one-step approach to multi-task visual grounding,” in Proc. NeurIPS , 2021
2021
Later among the works it cites.
Q. Huang, Y. Liang, J. Wei, Y. Cai, H. Liang, H.-f. Leung, and Q. Li, “Image difference captioning with instance-level fine-grained feature representation,” IEEE MultiMedia , vol. 24, pp. 2004–2017, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Chen, W. Ma, J. Xiao, H. Zhang, and S.-F. Chang, “Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,” in Proc. AAAI , vol. 35, no. 2, 2021, pp. 1036–1044
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
D. Liu, H. Zhang, F. Wu, and Z.-J. Zha, “Learning to assemble neural module tree networks for visual grounding,” in Proc. ICCV , 2019, pp. 4673–4682
2019
Cited alongside, same era.
R. Hong, D. Liu, X. Mo, X. He, and H. Zhang, “Learning to compose and reason with language tree structures for visual grounding,” IEEE Trans. Pattern Anal. Mach. Intell. , 2019
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Proc. NeurIPS , 2019
2019
Cited alongside, same era.
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proc. CVPR , 2019, pp. 658–666
2019
Cited alongside, same era.
G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji, “Multi-task collaborative network for joint referring expression comprehension and segmentation,” in Proc. CVPR , 2020, pp. 10 034–10 043
2020
Cited alongside, same era.
J. Liu, W. Wang, L. Wang, and M.-H. Yang, “Attribute-guided attention for referring expression generation and comprehension,” IEEE Trans. Image Process. , vol. 29, pp. 5244–5258, 2020
2020
Cited alongside, same era.
Y. Pan, T. Yao, Y. Li, and T. Mei, “X-linear attention networks for image captioning,” in Proc. CVPR , 2020, pp. 10 971–10 980
2020
Cited alongside, same era.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in Proc. ICCV , 2021, pp. 1780–1790
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. ICCV , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
H. Wang, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “Max-deeplab: End-to-end panoptic segmentation with mask transformers,” in Proc. CVPR , 2021, pp. 5463–5474
2021
Later among the works it cites.
B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” in Proc. NeurIPS , 2021
2021
Later among the works it cites.
X. Wang, Y. Chen, and W. Zhu, “A survey on curriculum learning,” IEEE Trans. Pattern Anal. Mach. Intell. , 2021
2021
Later among the works it cites.
X. Wu, J. Chang, Y.-K. Lai, J. Yang, and Q. Tian, “Bispl: Bidirectional self-paced learning for recognition from web data,” IEEE Trans. Image Process. , vol. 30, pp. 6512–6527, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Zhou, R. Ji, G. Luo, X. Sun, J. Su, X. Ding, C.-W. Lin, and Q. Tian, “A real-time global inference network for one-stage referring expression comprehension,” IEEE Trans. Neural Networks and Learning Systems. , 2021
2021
Later among the works it cites.
F. Yu, J. Tang, W. Yin, Y. Sun, H. Tian, H. Wu, and H. Wang, “Ernie-vil: Knowledge enhanced vision-language representations through scene graphs,” in Proc. AAAI , vol. 35, no. 4, 2021, pp. 3208–3216
2021
Later among the works it cites.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” in Proc. ICCV , 2015, pp. 19–27
2021
Later among the works it cites.
W. Jiang, M. Zhu, Y. Fang, G. Shi, X. Zhao, and Y. Liu, “Visual cluster grounding for image captioning,” IEEE Trans. Image Process. , 2022
2022
Closest in time.
Y. Liao, A. Zhang, Z. Chen, T. Hui, and S. Liu, “Progressive language-customized visual feature learning for one-stage visual grounding,” IEEE Trans. Image Process. , vol. 31, pp. 4266–4277, 2022
2022
Closest in time.
H. Zhao, J. T. Zhou, and Y.-S. Ong, “Word2pix: Word to pixel cross-attention transformer in visual grounding,” IEEE Trans. Neural Networks and Learning Systems. , 2022
2022
Closest in time.
M. Sun, W. Suo, P. Wang, Y. Zhang, and Q. Wu, “A proposal-free one-stage framework for referring expression comprehension and generation via dense cross-attention,” IEEE MultiMedia , 2022
2022
Closest in time.
2022
Closest in time.
W. Su, P. Miao, H. Dou, Y. Fu, and X. Li, “Referring expression comprehension using language adaptive inference,” in Proc. AAAI , vol. 37, no. 2, 2023, pp. 2357–2365
2023
Closest in time.
W. Su, P. Miao, H. Dou, G. Wang, L. Qiao, Z. Li, and X. Li, “Language adaptive weight generation for multi-task visual grounding,” in Proc. CVPR , 2023, pp. 10 857–10 866
2023
Closest in time.
S. Chen and Q. Zhao, “Divide and conquer: Answering questions with object factorization and compositional reasoning,” in Proc. CVPR , 2023, pp. 6736–6745
2023
Closest in time.
J. Zhu and H. Wang, “Multi-modal structure-embedding graph transformer for visual commonsense reasoning,” IEEE MultiMedia , 2023
2023
Closest in time.
Z. Li, Y. Guo, K. Wang, Y. Wei, L. Nie, and M. Kankanhalli, “Joint answering and explanation for visual commonsense reasoning,” IEEE Trans. Image Process. , 2023
2023
Closest in time.
Q. Yang, M. Ye, Z. Cai, K. Su, and B. Du, “Composed image retrieval via cross relation network with hierarchical aggregation transformer,” IEEE Trans. Image Process. , 2023
2023
Closest in time.