Fetching the paper…
Reading the bibliography…
Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A large-scale hierarchical image database. In: Proc. CVPR (2009)
2009
Earlier work this paper cites.
Ordonez, V., Kulkarni, G., Berg, T.: Im2text: Describing images using 1 million captioned photographs. In: Proc. NIPS (2011)
2011
Earlier work this paper cites.
Kazemzadeh, S., Ordonez, V., Matten, M.A., Berg, T.L.: ReferItGame: Referring to Objects in Photographs of Natural Scenes. In: Proc. EMNLP (2014)
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common Objects in Context. In: Proc. ECCV (2014)
2014
Earlier work this paper cites.
Fang, H., Gupta, S., Iandola, F., Srivastava, R.K., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J.C., et al.: From Captions to Visual Concepts and Back. In: Proc. CVPR (2015)
2015
Earlier work this paper cites.
Karpathy, A., Fei-Fei, L.: Deep Visual-semantic Alignments for Generating Image Descriptions. In: Proc. CVPR (2015)
2015
Earlier work this paper cites.
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. In: Proc. ICCV (2015)
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks. In: Proc. NIPS (2015)
2015
Earlier work this paper cites.
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. In: Proc. EMNLP (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 770–778 (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Johnson, J., Karpathy, A., Fei-Fei, L.: DenseCap: Fully Convolutional Localization Networks for Dense Captioning. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M.S., Fei-Fei, L.: Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Computer Vision
2016
Earlier work this paper cites.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S.E., Fu, C., Berg, A.C.: SSD: Single Shot MultiBox Detector. In: Proc. ECCV (2016)
2016
Earlier work this paper cites.
Mao, J., Huang, J., Toshev, A., Camburu, O.M., Yuille, A.L., Murphy, K.P.: Generation and Comprehension of Unambiguous Object Descriptions. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Redmon, J., Divvala, S.K., Girshick, R.B., Farhadi, A.: You Only Look Once: Unified, Real-Time Object Detection. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Yu, L., Poirson, P., Yang, S., Berg, A.C., Berg, T.L.: Modeling Context in Referring Expressions. In: Proc. ECCV (2016)
2016
Earlier work this paper cites.
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T.A., Nießner, M.: ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. In: Proc. CVPR (2017)
2017
Earlier work this paper cites.
He, K., Gkioxari, G., Dollár, P., Girshick, R.B.: Mask R-CNN. In: Proc. ICCV (2017)
2017
Earlier work this paper cites.
Hu, R., Rohrbach, M., Andreas, J., Darrell, T., Saenko, K.: Modeling Relationships in Referential Expressions with Compositional Modular Networks. In: Proc. CVPR (2017)
2017
Cited alongside, same era.
Lin, T.Y., Goyal, P., Girshick, R.B., He, K., Dollár, P.: Focal Loss for Dense Object Detection. In: Proc. ICCV (2017)
2017
Cited alongside, same era.
Qi, C., Yi, L., Su, H., Guibas, L.J.: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. In: Proc. NIPS (2017)
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is All you Need. In: Proc. NIPS (2017)
2017
Cited alongside, same era.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In: Proc. CVPR (2018)
Gan, Z., Chen, Y.C., Li, L., Zhu, C., Cheng, Y., Liu, J.: Large-Scale Adversarial Training for Vision-and-Language Representation Learning. In: Proc. NeurIPS (2020)
2020
Later among the works it cites.
Lu, J., Goswami, V., Rohrbach, M., Parikh, D., Lee, S.: 12-in-1: Multi-Task Vision and Language Representation Learning. In: Proc. CVPR (2020)
2020
Later among the works it cites.
Yang, Z., Chen, T., Wang, L., Luo, J.: Improving one-stage visual grounding by recursive sub-query construction. In: Proc. ECCV (2020)
2020
Later among the works it cites.
Deng, J., Yang, Z., Chen, T., Zhou, W., Li, H.: Transvg: End-to-end visual grounding with transformers. In: Proc. ICCV (2021)
2021
Closest in time.
Feng, M., Li, Z., Li, Q., Zhang, L., Zhang, X., Zhu, G., Zhang, H., Wang, Y., Mian, A.: Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud. In: Proc. ICCV (2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning. In: Proc. ACL (2018)
2018
Cited alongside, same era.
Yu, Z., Yu, J., Xiang, C., Zhao, Z., Tian, Q., Tao, D.: Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding. In: Proc. IJCAI (2018)
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Lu, J., Batra, D., Parikh, D., Lee, S.: ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In: Proc. NeurIPS (2019)
2019
Cited alongside, same era.
Qi, C., Litany, O., He, K., Guibas, L.J.: Deep Hough Voting for 3D Object Detection in Point Clouds. In: Proc. ICCV (2019)
2019
Cited alongside, same era.
2021
Closest in time.
He, D., Zhao, Y., Luo, J., Hui, T., Huang, S., Zhang, A., Liu, S.: TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding. In: Proc. ACMMM (2021)
2021
Closest in time.
Huang, P.H., Lee, H.H., Chen, H.T., Liu, T.L.: Text-Guided Graph Neural Networks for Referring 3D Instance Segmentation. In: Proc. AAAI (2021)
2021
Closest in time.
Jaegle, A., Gimeno, F., Brock, A., Zisserman, A., Vinyals, O., Carreira, J.: Perceiver: General Perception with Iterative Attention. In: Proc. ICML (2021)
2021
Closest in time.
Kamath, A., Singh, M., LeCun, Y.A., Misra, I., Synnaeve, G., Carion, N.: MDETR - Modulated Detection for End-to-End Multi-Modal Understanding. In: Proc. ICCV (2021)
2021
Closest in time.
Liu, Z., Zhang, Z., Cao, Y., Hu, H., Tong, X.: Group-Free 3D Object Detection via Transformers. In: Proc. ICCV (2021)
2021
Closest in time.
Misra, I., Girdhar, R., Joulin, A.: An End-to-End Transformer Model for 3D Object Detection. In: Proc. ICCV (2021)
2021
Closest in time.
Roh, J., Desingh, K., Farhadi, A., Fox, D.: LanguageRefer: Spatial-Language Model for 3D Visual Grounding. In: Proc. CoRL (2021)
2021
Closest in time.
Yang, Z., Zhang, S., Wang, L., Luo, J.: SAT: 2D Semantics Assisted Training for 3D Visual Grounding. In: Proc. ICCV (2021)
2021
Closest in time.
Yuan, Z., Yan, X., Liao, Y., Zhang, R., Li, Z., Cui, S.: InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring. In: Proc. ICCV (2021)
2021
Closest in time.
Zhao, L., Cai, D., Sheng, L., Xu, D.: 3DVG-Transformer: Relation Modeling for Visual Grounding on Point Clouds. In: Proc. ICCV (2021)
2021
Closest in time.
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable DETR: Deformable Transformers for End-to-End Object Detection. In: Proc. ICLR (2021)
2021
Closest in time.
Abdelreheem, A., Upadhyay, U., Skorokhodov, I., Yahya, R.A., Chen, J., Elhoseiny, M.: 3DRefTransformer: Fine-Grained Object Identification in Real-World Scenes Using Natural Language. In: Proc. WACV (2022)
2022
Closest in time.
Li, L.H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.N., Chang, K.W., Gao, J.: Grounded Language-Image Pre-training. In: Proc. CVPR (2022)
2022
Closest in time.