Fetching the paper…
Reading the bibliography…
Referring expression comprehension (REC) involves localizing a target instance based on a textual description.
Maf: Multimodal alignment framework for weakly-supervised phrase grounding
Q. Wang, H. Tan, S. Shen, M. W. Mahoney, and Z. Yao · 2010
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
H. De Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Referring expression generation and comprehension via attributes
J. Liu, L. Wang, and M.-H. Yang · 2017
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Earlier work this paper cites.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
R. Liu, C. Liu, Y. Bai, and A. L. Yuille · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2019
Earlier work this paper cites.
Referring expression comprehension with semantic visual relationship and word mapping
C. Zhang, W. Li, W. Ouyang, Q. Wang, W.-S. Kim, and S. Hong · 2019
Earlier work this paper cites.
Reasoning visual dialogs with structural and partial observations
Z. Zheng, W. Wang, S. Qi, and S.-C. Zhu · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
Cops-ref: A new dataset and task on compositional referring expression comprehension
Z. Chen, P. Wang, L. Ma, K.-Y. K. Wong, and Q. Wu · 2020
Earlier work this paper cites.
Refer360 degree: A referring expression recognition dataset in 360 degree images
V. Cirik, T. Berg-Kirkpatrick, and L.-P. Morency · 2020
Cited alongside, same era.
Referring expression comprehension: A survey of methods and datasets
Y. Qiao, C. Deng, and Q. Wu · 2020
Cited alongside, same era.
Phrasecut: Language-based image segmentation in the wild
C. Wu, Z. Lin, S. Cohen, T. Bui, and S. Maji · 2020
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Cited alongside, same era.
Scene-text oriented referring expression comprehension
Y. Bu, L. Li, J. Xie, Q. Liu, Y. Cai, Q. Huang, and Q. Li · 2022
Cited alongside, same era.
Glamm: Pixel grounding large multimodal model
H. Rasheed, M. Maaz, S. Shaji, A. Shaker, S. Khan, H. Cholakkal, R. M. Anwer, E. Xing, M.-H. Yang, and F. S. Khan · 2023
Later among the works it cites.
Pixellm: Pixel reasoning with large multimodal model
Z. Ren, Z. Huang, Y. Wei, Y. Zhao, D. Fu, J. Feng, and X. Jin · 2023
Later among the works it cites.
Lenna: Language enhanced reasoning detection assistant
F. Wei, X. Zhang, A. Zhang, B. Zhang, and X. Chu · 2023
Later among the works it cites.
Universal instance perception as object discovery and retrieval
B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Hao, H. Song, L. Dong, S. Huang, Z. Chi, W. Wang, S. Ma, and F. Wei · 2022
Cited alongside, same era.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2022
Cited alongside, same era.
Refcrowd: Grounding the target in crowd with referring expressions
H. Qiu, H. Li, T. Zhao, L. Wang, Q. Wu, and F. Meng · 2022
Cited alongside, same era.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Cited alongside, same era.
Glipv2: Unifying localization and vision-language understanding
H. Zhang, P. Zhang, X. Hu, Y.-C. Chen, L. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao · 2022
Cited alongside, same era.
Towards unifying reference expression generation and comprehension
D. Zheng, T. Kong, Y. Jing, J. Wang, and X. Wang · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Ferret: Refer and ground anything anywhere at any granularity
H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang · 2023
Later among the works it cites.
Griffon: Spelling out all object locations at any granularity with large language models
Y. Zhan, Y. Zhu, Z. Chen, F. Yang, M. Tang, and J. Wang · 2023
Later among the works it cites.
Llava-grounding: Grounded visual chat with large multimodal models
H. Zhang, H. Li, F. Li, T. Ren, X. Zou, S. Liu, S. Huang, J. Gao, L. Zhang, C. Li, et al · 2023
Later among the works it cites.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
P. Gao, R. Zhang, C. Liu, L. Qiu, S. Huang, W. Lin, S. Zhao, S. Geng, Z. Lin, P. Jin, et al · 2024
Closest in time.
Multi-modal instruction tuned llms with fine-grained visual perception
J. He, Y. Wang, L. Wang, H. Lu, J.-Y. He, J.-P. Lan, B. Luo, and X. Xie · 2024
Closest in time.
Relationvlm: Making large vision-language models understand visual relations
Z. Huang, Z. Zhang, Z.-J. Zha, Y. Lu, and B. Guo · 2024
Closest in time.
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
B. Jia, Y. Chen, H. Yu, Y. Wang, X. Niu, T. Liu, Q. Li, and S. Huang · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Lego: Language enhanced multi-modal grounding model
Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. T. Vu, et al · 2024
Closest in time.
Groma: Localized visual tokenization for grounding multimodal large language models
C. Ma, Y. Jiang, J. Wu, Z. Yuan, and X. Qi · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
A. Meta · 2024
Closest in time.
Generalizable entity grounding via assistance of large language model
L. Qi, Y.-W. Chen, L. Yang, T. Shen, X. Li, W. Guo, Y. Xu, and M.-H. Yang · 2024
Closest in time.
Groundvlp: Harnessing zero-shot visual grounding from vision-language pre-training and open-vocabulary object detection
H. Shen, T. Zhao, M. Zhu, and J. Yin · 2024
Closest in time.
Y. Zhan, Y. Zhu, H. Zhao, F. Yang, M. Tang, and J. Wang · 2024
Closest in time.
Llm-optic: Unveiling the capabilities of large language models for universal visual grounding
H. Zhao, W. Ge, and Y.-c. Chen · 2024
Closest in time.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee · 2024
Closest in time.