Fetching the paper…
Reading the bibliography…
Highlighting particularly relevant regions of an image can improve the performance of vision-language models (VLMs) on various vision-language (VL) tasks by guiding the model to attend more closely to these regions of interest.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O.-M. Camburu, A. L. Yuille, and K. P. Murphy · 2015
Earlier work this paper cites.
“why should I trust you?”: Explaining the predictions of any classifier
M. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2017
Earlier work this paper cites.
Taking a hint: Leveraging explanations to make vision and language models more grounded
R. R. Selvaraju, S. Lee, Y. Shen, H. Jin, S. Ghosh, L. Heck, D. Batra, and D. Parikh · 2019
Earlier work this paper cites.
Self-critical reasoning for robust visual question answering
J. Wu and R. Mooney · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
spacy: Industrial-strength natural language processing in python, 2020
M. Honnibal, I. Montani, S. V. Landeghem, and A. Boyd · 2020
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2021
Earlier work this paper cites.
Mdetr–modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models, 2021
Y. Yao, A. Zhang, Z. Zhang, Z. Liu, T.-S. Chua, and M. Sun · 2021
Earlier work this paper cites.
MERLOT: Multimodal Neural Script Knowledge Models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Earlier work this paper cites.
Exploring visual prompts for adapting large-scale models, 2022
H. Bahng, A. Jahanian, S. Sankaranarayanan, and P. Isola · 2022
Cited alongside, same era.
Visual Prompting via Image Inpainting
A. Bar, Y. Gandelsman, T. Darrell, A. Globerson, and A. A. Efros · 2022
Cited alongside, same era.
Focalclick: Towards practical interactive image segmentation
X. Chen, Z. Zhao, Y. Zhang, M. Duan, D. Qi, and H. Zhao · 2022
Cited alongside, same era.
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Cited alongside, same era.
Grounded language-image pre-training
L. H. Li*, P. Zhang*, H. Zhang*, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Cited alongside, same era.
Answer questions with right image regions: A visual attention regularization approach
Guiding Image Captioning Models Toward More Specific Captions
S. Kornblith, L. Li, Z. Wang, and T. Nguyen · 2023
Later among the works it cites.
S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, and L. Bing · 2023
Later among the works it cites.
Crepe: Can vision-language foundation models reason compositionally?
Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna · 2023
Later among the works it cites.
Contrastive decoding improves reasoning in large language models, 2023
S. O’Brien and M. Lewis · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI, :, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Łukasz Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Łukasz Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. Mély, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. O’Keefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Liu, Y. Guo, J. Yin, X. Song, W. Liu, L. Nie, and M. Zhang · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, R. Gontijo-Lopes, B. K. Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Cited alongside, same era.
FLAVA: A foundational language and vision alignment model
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela · 2022
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Cited alongside, same era.
Visfis: Visual feature importance supervision with right-for-the-right-reason objectives
Z. Ying, P. Hase, and M. Bansal · 2022
Cited alongside, same era.
Glipv2: Unifying localization and vision-language understanding
H. Zhang, P. Zhang, X. Hu, Y.-C. Chen, L. H. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao · 2022
Cited alongside, same era.
Making Large Multimodal Models Understand Arbitrary Visual Prompts, 2023
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y. Chai, D. Park, and Y. J. Lee · 2023
Cited alongside, same era.
Later among the works it cites.
Stay on topic with classifier-free guidance
G. Sanchez, H. Fan, A. Spangher, E. Levi, P. S. Ammanamanchi, and S. Biderman · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding, 2023
W. Shi, X. Han, M. Lewis, Y. Tsvetkov, L. Zettlemoyer, and S. W. tau Yih · 2023
Later among the works it cites.
What does clip know about a red circle? visual prompt engineering for vlms
A. Shtedritski, C. Rupprecht, and A. Vedaldi · 2023
Later among the works it cites.
Alpha-clip: A clip model focusing on wherever you want, 2023
Z. Sun, Y. Fang, T. Wu, P. Zhang, Y. Zang, S. Kong, Y. Xiong, D. Lin, and J. Wang · 2023
Later among the works it cites.
Imagen editor and editbench: Advancing and evaluating text-guided image inpainting
S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pellegrini, Y. Onoe, S. Laszlo, D. J. Fleet, R. Soricut, J. Baldridge, M. Norouzi, P. Anderson, and W. Chan · 2023
Later among the works it cites.
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V, 2023
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Later among the works it cites.
What you see is what you read? improving text-image alignment evaluation
M. Yarom, Y. Bitton, S. Changpinyo, R. Aharoni, J. Herzig, O. Lang, E. Ofek, and I. Szpektor · 2023
Later among the works it cites.
Gpt4roi: Instruction tuning large language model on region-of-interest, 2023
S. Zhang, P. Sun, S. Chen, M. Xiao, W. Shao, W. Zhang, Y. Liu, K. Chen, and P. Luo · 2023
Later among the works it cites.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee · 2023
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
cola: A benchmark for compositional text-to-image retrieval
A. Ray, F. Radenovic, A. Dubey, B. Plummer, R. Krishna, and K. Saenko · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang · 2024
Closest in time.
Mitigating object hallucination in large vision-language models via classifier-free guidance, 2024
L. Zhao, Y. Deng, W. Zhang, and Q. Gu · 2024
Closest in time.