Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG).
Zhou, X.; Wang, D.; and Krähenbühl, P. 2019 · 1904
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019 · 1910
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. L. 2011 · 2011
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A.; and Fei-Fei, L. 2014 · 2014
Earlier work this paper cites.
Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
Karpathy, A.; Joulin, A.; and Fei-Fei, L. 2014 · 2014
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. L. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Natural Language Object Retrieval
Hu, R.; Xu, H.; Rohrbach, M.; Feng, J.; Saenko, K.; and Darrell, T. 2015 · 2015
Earlier work this paper cites.
Generation and Comprehension of Unambiguous Object Descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.-M.; Yuille, A. L.; and Murphy, K. 2015 · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Learning Deep Structure-Preserving Image-Text Embeddings
Wang, L.; Li, Y.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Fukui, A.; Park, D. H.; Yang, D.; Rohrbach, A.; Darrell, T.; and Rohrbach, M. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Modeling Relationships in Referential Expressions with Compositional Modular Networks
Hu, R.; Rohrbach, M.; Andreas, J.; Darrell, T.; and Saenko, K. 2016 · 2016
Earlier work this paper cites.
Modeling Context Between Objects for Referring Expression Understanding
Nagaraja, V. K.; Morariu, V. I.; and Davis, L. S. 2016 · 2016
Earlier work this paper cites.
CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples
Radenovic, F.; Tolias, G.; and Chum, O. 2016 · 2016
Earlier work this paper cites.
YFCC100M: the new data in multimedia research
Thomee, B.; Shamma, D. A.; Friedland, G.; Elizalde, B.; Ni, K.; Poland, D. N.; Borth, D.; and Li, L.-J. 2016 · 2016
Earlier work this paper cites.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2017 · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M.; and Montani, I. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Referring Expression Generation and Comprehension via Attributes
Liu, J.; Wang, L.; and Yang, M.-H. 2017 · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Earlier work this paper cites.
Conditional Image-Text Embedding Networks
Plummer, B. A.; Kordas, P.; Kiapour, M. H.; Zheng, S.; Piramuthu, R.; and Lazebnik, S. 2017 · 2017
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Learning Two-Branch Neural Networks for Image-Text Matching Tasks
Wang, L.; Li, Y.; Huang, J.; and Lazebnik, S. 2017 · 2017
Cited alongside, same era.
Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding
Akbari, H.; Karaman, S.; Bhargava, S.; Chen, B.; Vondrick, C.; and Chang, S.-F. 2018 · 2018
Cited alongside, same era.
Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
Chen, X.; Ma, L.; Chen, J.; Jie, Z.; Liu, W.; and Luo, J. 2018 · 2018
Cited alongside, same era.
Bilinear Attention Networks
Kim, J.-H.; Jun, J.; and Zhang, B.-T. 2018 · 2018
Cited alongside, same era.
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Dynamic DETR: End-to-End Object Detection With Dynamic Attention
Dai, X.; Chen, Y.; Yang, J.; Zhang, P.; Yuan, L.; and Zhang, L. 2021 · 2021
Later among the works it cites.
TransVG: End-to-End Visual Grounding with Transformers
Deng, J.; Yang, Z.; Chen, T.; Zhou, W.; and Li, H. 2021 · 2021
Later among the works it cites.
Visual Grounding with Transformers
Du, Y.; Fu, Z.; Liu, Q.; and Wang, Y. 2021 · 2021
Later among the works it cites.
CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
Fürst, A.; Rumetshofer, E.; Tran, V. H.; Ramsauer, H.; Tang, F.; Lehner, J. M.; Kreil, D. P.; Kopp, M. K.; Klambauer, G.; Bitto-Nemling, A.; and Hochreiter, S. 2021 · 2021
Later among the works it cites.
Fast convergence of detr with spatially modulated co-attention
Gao, P.; Zheng, M.; Wang, X.; Dai, J.; and Li, H. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; Duerig, T.; and Ferrari, V. 2018 · 2018
Cited alongside, same era.
Learning to Assemble Neural Module Tree Networks for Visual Grounding
Liu, D.; Zhang, H.; Wu, F.; and Zha, Z.-J. 2018 · 2018
Cited alongside, same era.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Cited alongside, same era.
YOLOv3: An Incremental Improvement
Redmon, J.; and Farhadi, A. 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
G3raphGround: Graph-Based Language Grounding
Bajaj, M.; Wang, L.; and Sigal, L. 2019 · 2019
Cited alongside, same era.
Neural Sequential Phrase Grounding (SeqGROUND)
Dogan, P.; Sigal, L.; and Gross, M. 2019 · 2019
Cited alongside, same era.
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Later among the works it cites.
Grounded Language-Image Pre-training
Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. 2021 · 2021
Later among the works it cites.
Referring Transformer: A One-step Approach to Multi-task Visual Grounding
Li, M.; and Sigal, L. 2021 · 2021
Later among the works it cites.
Conditional DETR for Fast Training Convergence
Meng, D.; Chen, X.; Fan, Z.; Zeng, G.; Li, H.; Yuan, Y.; Sun, L.; and Wang, J. 2021 · 2021
Later among the works it cites.
Disentangled Motif-aware Graph Learning for Phrase Grounding
Mu, Z.; Tang, S.; Tan, J.; Yu, Q.; and Zhuang, Y. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Anchor DETR: Query Design for Transformer-Based Detector
Wang, Y.; Zhang, X.; Yang, T.; and Sun, J. 2021 · 2021
Later among the works it cites.
VinVL: Revisiting Visual Representations in Vision-Language Models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2021 · 2021
Later among the works it cites.
Mask2Former for Video Instance Segmentation
Cheng, B.; Choudhuri, A.; Misra, I.; Kirillov, A.; Girdhar, R.; and Schwing, A. G. 2022 · 2022
Closest in time.
Deconfounded Visual Grounding
Huang, J.; Qin, Y.; Qi, J.; Sun, Q.; and Zhang, H. 2022 · 2022
Closest in time.
DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L. M.; and Zhang, L. 2022 · 2022
Closest in time.
Adapting CLIP For Phrase Localization Without Further Training
Li, J.; Shakhnarovich, G.; and Yeh, R. 2022 · 2022
Closest in time.
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; and Zhang, L. 2022 · 2022
Closest in time.
Referring Expression Comprehension via Cross-Level Multi-Modal Fusion
Miao, P.; Su, W.; Wang, L.; Fu, Y.; and Li, X. 2022 · 2022
Closest in time.
Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Wang, P.; Yang, A.; Men, R.; Lin, J.; Bai, S.; Li, Z.; Ma, J.; Zhou, C.; Zhou, J.; Yang, H.; and Zhou, C. 2022 · 2022
Closest in time.
Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning
Yang, L.; Xu, Y.; Yuan, C.; Liu, W.; Li, B.; and Hu, W. 2022 · 2022
Closest in time.
Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding
Ye, J.; Tian, J.; Yan, M.; Yang, X.; Wang, X.; Zhang, J.; He, L.; and Lin, X. 2022 · 2022
Closest in time.
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L. M.; and Shum, H.-Y. 2022 · 2022
Closest in time.
SeqTR: A Simple yet Universal Network for Visual Grounding
Zhu, C.; Zhou, Y.; Shen, Y.; Luo, G.; Pan, X.; Lin, M.; Chen, C.; Cao, L.; Sun, X.; and Ji, R. 2022 · 2022
Closest in time.