An image is often considered worth a thousand words, and certain images can tell rich and insightful stories.
Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cultural signs, and symbols, are vital to every culture.
Our research compares the performance of four popular vision-language models (GPT-4V, Gemini Pro Vision, LLaVA, and OpenFlamingo) in identifying culturally specific information in such images and creating accurate and culturally sensitive image captions.
We also propose a new evaluation metric, the Cultural Awareness Score (CAS), which measures the degree of cultural awareness in image captions.
How Culturally Aware are Vision-Language Models? · Around
Built on
S. Sheng, A. N. Venkitasubramanian, and M.-F. Moens, “A Markov network based passage retrieval method for multimodal question answering in the cultural heritage domain,” in Lecture notes in computer science, 2018, pp. 3–15
O. Burda-Lassen, “Ukrainian-To-English Folktale Corpus: Parallel Corpus Creation and augmentation for Machine Translation in Low-Resource Languages,” ACL Anthology, Sep. 01, 2022. https://aclanthology.org/2022.amta-coco4mt.4/
B. Zheng et al., “Image captioning for cultural artworks: a case study on ceramics,” Multimedia Systems, vol. 29, no. 6, pp. 3223–3243, Sep. 2023, doi: 10.1007/s00530-023-01178-8
J. Kharchenko, T. Roosta, A. Chadha, and C. Shah, “How well do LLMs represent values across cultures? Empirical analysis of LLM responses based on Hofstede Cultural dimensions,” arXiv.org, Jun. 21, 2024. https://arxiv.org/abs/2406.14805https://arxiv.org/abs/2302.14045