Fetching the paper…
Reading the bibliography…
Image Captioning generates descriptive sentences from images using Vision-Language Pre-trained models (VLPs) such as BLIP, which has improved greatly.
The dynamic elements of culture
Ben Halpern · 1955
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Demographics of mechanical turk
Panagiotis G Ipeirotis · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Weaqa: Weak supervision via captions for visual question answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral · 2020
Earlier work this paper cites.
Compare and reweight: Distinctive image captioning using similar images sets
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, and Antoni B Chan · 2020
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Visually grounded reasoning across languages and cultures
Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott · 2021
Cited alongside, same era.
Dataset diversity: measuring and mitigating geographical bias in image search and retrieval
Abhishek Mandal, Susan Leavy, and Suzanne Little · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford et al · 2021
Cited alongside, same era.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
William A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman · 2022
Later among the works it cites.
Revise: A tool for measuring and mitigating bias in visual datasets
Angelina Wang, Alexander Liu, Ryan Zhang, Anat Kleiman, Leslie Kim, Dora Zhao, Iroha Shirai, Arvind Narayanan, and Olga Russakovsky · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2022
Later among the works it cites.
Geomlama: Geo-diverse commonsense probing on multilingual pre-trained language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Cited alongside, same era.
Broaden the vision: Geo-diverse visual commonsense reasoning
Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, and Kai-Wei Chang · 2021
Cited alongside, same era.
Language bias in visual question answering: A survey and taxonomy
Desen Yuan · 2021
Cited alongside, same era.
All you may need for vqa are image captions
Soravit Changpinyo, Doron Kukliansky, Idan Szpektor, Xi Chen, Nan Ding, and Radu Soricut · 2022
Cited alongside, same era.
From images to textual prompts: Zero-shot vqa with frozen large language models
Jiaxian Guo, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Boyang Li, Dacheng Tao, and Steven CH Hoi · 2022
Cited alongside, same era.
Challenges and strategies in cross-cultural nlp
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al · 2022
Cited alongside, same era.
Vision-and-language pretrained models: A survey
Siqu Long, Feiqi Cao, Soyeon Caren Han, and Haiqin Yang · 2022
Cited alongside, same era.
Da Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li, and Kai-Wei Chang · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Cultural concept adaptation on multimodal reasoning
Zhi Li and Yin Zhang · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Extracting cultural commonsense knowledge at scale
Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum · 2023
Later among the works it cites.
Givl: Improving geographical inclusivity of vision-language models with pre-training methods
Da Yin, Feng Gao, Govind Thattai, Michael Johnston, and Kai-Wei Chang · 2023
Later among the works it cites.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions
Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny · 2023
Later among the works it cites.