Fetching the paper…
Reading the bibliography…
In this work, we explore neat yet effective Transformer-based frameworks for visual grounding.
X. Liu, Z. Wang, J. Shao, X. Wang, and H. Li, “Improving referring expression grounding with cross-modal attention-guided erasing,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 1950–1959
1959
Earlier work this paper cites.
P. Wang, Q. Wu, J. Cao, C. Shen, L. Gao, and A. v. d. Hengel, “Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 1960–1968
1968
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y. M. Lui, “Visual object tracking using adaptive correlation filters,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2010, pp. 2544–2550
2010
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur, “Recurrent neural network based language model,” in InterSpeech , 2010
2010
Earlier work this paper cites.
H. J. Escalante, C. A. Hernández, J. A. Gonzalez, A. López-López, M. Montes, E. F. Morales, L. E. Sucar, L. Villaseñor, and M. Grubinger, “The segmented and annotated iapr tc-12 benchmark,” Computer Vision and Image Understanding (CVIU) , vol. 114, pp. 419–428, 2010
2010
Earlier work this paper cites.
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,” International Journal of Computer Vision (IJCV) , vol. 104, pp. 154–171, 2013
2013
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2014, pp. 580–587
2014
Earlier work this paper cites.
J. F. Henriques, R. Caseiro, P. Martins, and J. Batista, “High-speed tracking with kernelized correlation filters,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 37, pp. 583–596, 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Annual Meeting of the Association for Computational Linguistics (ACL) , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2014, pp. 740–755
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1440–1448
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. K. Nagaraja, V. I. Morariu, and L. S. Davis, “Modeling context between objects for referring expression understanding,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2016, pp. 792–807
2016
Earlier work this paper cites.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural language object retrieval,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 4555–4564
2016
Earlier work this paper cites.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 5005–5013
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 11–20
2016
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks.” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv:1607.06450 , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Li, Y. Wei, X. Liang, F. Zhao, J. Li, T. Xu, and J. Feng, “Deep attribute-preserving metric learning for natural language object retrieval,” in Proceedings of the 28th ACM International Conference on Multimedia (ACM MM) . ACM, 2017, pp. 181–189
2017
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” International Journal of Computer Vision (IJCV) , vol. 123, no. 1, p. 74, 2017
2017
Earlier work this paper cites.
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko, “Modeling relationships in referential expressions with compositional modular networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 1115–1124
2017
Earlier work this paper cites.
Y. Zhang, L. Yuan, Y. Guo, Z. He, I.-A. Huang, and H. Lee, “Discriminative bimodal networks for visual localization and detection with natural language queries,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 557–566
2017
Earlier work this paper cites.
K. Chen, R. Kovvuri, and R. Nevatia, “Query-guided regression network with context policy for phrase grounding,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 824–832
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 2961–2969
2017
Cited alongside, same era.
L. Wang, Y. Li, J. Huang, and S. Lazebnik, “Learning two-branch neural networks for image-text matching tasks,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 41, pp. 394–407, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 1307–1315
2018
Cited alongside, same era.
——, “Graph-structured referring expression reasoning in the wild,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 9952–9961
2020
Later among the works it cites.
S. Yang, G. Li, and Y. Yu, “Relationship-embedded representation learning for grounding referring expressions,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 43, no. 8, pp. 2765–2779, 2020
2020
Later among the works it cites.
S. Yang, G. Li, and Y. Yu, “Propagating over phrase relations for one-stage visual grounding,” in Proceedings of the European Conference on Computer Vision (ECCV) . Springer, 2020, pp. 589–605
2020
Later among the works it cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2020, pp. 213–229
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhang, Y. Niu, and S.-F. Chang, “Grounding referring expressions in images by variational context,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4158–4166
2018
Cited alongside, same era.
B. Zhuang, Q. Wu, C. Shen, I. Reid, and A. van den Hengel, “Parallel attention: A unified framework for visual object discovery through dialogs and queries,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4252–4261
2018
Cited alongside, same era.
B. A. Plummer, P. Kordas, M. H. Kiapour, S. Zheng, R. Piramuthu, and S. Lazebnik, “Conditional image-text embedding networks,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 249–264
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2018, pp. 6077–6086
2018
Cited alongside, same era.
J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv:1804.02767 , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International Conference on Machine Learning (ICML) , 2018, pp. 4055–4064
2018
Cited alongside, same era.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser, “Universal transformers,” in International Conference on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in International Conference on Machine Learning (ICML) , 2020, pp. 1691–1703
2020
Later among the works it cites.
F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo, “Learning texture transformer network for image super-resolution,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 5791–5800
2020
Later among the works it cites.
Y. Zeng, J. Fu, and H. Chao, “Learning joint spatial-temporal transformations for video inpainting,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2020, pp. 528–543
2020
Later among the works it cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2020, pp. 104–120
2020
Later among the works it cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2020, pp. 121–137
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Liu, B. Wan, X. Zhu, and X. He, “Learning cross-modal context graph for visual grounding,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , vol. 34, no. 07, 2020, pp. 11 645–11 652
2020
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
J. Deng, Z. Yang, T. Chen, W. Zhou, and H. Li, “Transvg: End-to-end visual grounding with transformers,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2021, pp. 1769–1779
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML) , 2021, pp. 10 347–10 357
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
D. Meng, X. Chen, Z. Fan, G. Zeng, H. Li, Y. Yuan, L. Sun, and J. Wang, “Conditional detr for fast training convergence,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2021, pp. 3651–3660
2021
Later among the works it cites.
2021
Later among the works it cites.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
J. Li, R. R. Selvaraju, A. D. Gotmare, S. Joty, C. Xiong, and S. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2021, pp. 1780–1790
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Li and L. Sigal, “Referring transformer: A one-step approach to multi-task visual grounding,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
Y. Du, Z. Fu, Q. Liu, and Y. Wang, “Visual grounding with transformers,” in Proceedings of the IEEE International Conference on Multimedia & Expo (ICME) , 2022
2022
Closest in time.