Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge present in both visual and textual data.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
A unified framework for multilingual and code-mixed visual question answering
D. Gupta, P. Lenka, A. Ekbal, and P. Bhattacharyya · 2020
Earlier work this paper cites.
Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages, 2020
A. Kunchukuttan, D. Kakwani, S. Golla, G. N. C., A. Bhattacharyya, M. M. Khapra, and P. Kumar · 2020
Earlier work this paper cites.
Masakhane – machine translation for africa, 2020
I. Orife, J. Kreutzer, B. Sibanda, D. Whitenack, K. Siminyu, L. Martinus, J. T. Ali, J. Abbott, V. Marivate, S. Kabongo, M. Meressa, E. Murhabazi, O. Ahia, E. van Biljon, A. Ramkilowan, A. Akinfaderin, A. Öktem, W. Akin, G. Kioko, K. Degila, H. Kamper, B. Dossou, C. Emezue, K. Ogueji, and A. Bashir · 2020
Earlier work this paper cites.
A framework for few-shot language model evaluation, 2021
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Visually grounded reasoning across languages and cultures
F. Liu, E. Bugliarello, E. M. Ponti, S. Reddy, N. Collier, and D. Elliott · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
A survey of methods, datasets and evaluation metrics for visual question answering
H. Sharma and A. S. Jalal · 2021
Cited alongside, same era.
SUTD-TrafficQA: A question answering benchmark and an efficient network for video reasoning over traffic events
L. Xu, H. Huang, and J. Liu · 2021
Cited alongside, same era.
Just ask: Learning to answer questions from millions of narrated videos
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2021
Cited alongside, same era.
Cross-lingual and multilingual CLIP
Bhasa: A holistic southeast asian linguistic and cultural evaluation suite for large language models, 2023
W. Q. Leong, J. G. Ngui, Y. Susanto, H. Rengarajan, K. Sarveswaran, and W. C. Tjhi · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Later among the works it cites.
MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Later among the works it cites.
Global voices, local biases: Socio-cultural prejudices across languages
A. Mukherjee, C. Raj, Z. Zhu, and A. Anastasopoulos · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Carlsson, P. Eisen, F. Rekathati, and M. Sahlgren · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
InfographicVQA
M. Mathew, V. Bagal, R. P. Tito, D. Karatzas, E. Valveny, and C. V. Jawahar · 2022
Cited alongside, same era.
xGQA: Cross-lingual visual question answering
J. Pfeiffer, G. Geigle, A. Kamath, J.-M. Steitz, S. Roth, I. Vulić, and I. Gurevych · 2022
Cited alongside, same era.
Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, et al · 2023
Cited alongside, same era.
MaXM: Towards multilingual visual question answering
S. Changpinyo, L. Xue, M. Yarom, A. Thapliyal, I. Szpektor, J. Amelot, X. Chen, and R. Soricut · 2023
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
P. Pezeshkpour and E. Hruschka · 2023
Later among the works it cites.
Leveraging large language models for multiple choice question answering
J. Robinson and D. Wingate · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
G. Team · 2023
Later among the works it cites.
Copal-id: Indonesian language reasoning with local culture and nuances
H. A. Wibowo, E. H. Fuadi, M. N. Nityasya, R. E. Prasojo, and A. F. Aji · 2023
Later among the works it cites.
MM-Vet: Evaluating large multimodal models for integrated capabilities
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Later among the works it cites.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Towards measuring and modeling "culture" in LLMs: A survey
M. F. Adilazuarda, S. Mukherjee, P. Lavania, S. Singh, A. Dwivedi, A. F. Aji, J. O’Neill, A. Modi, and M. Choudhury · 2024
Closest in time.
Multimodal foundation models: From specialists to general-purpose assistants
C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, and J. Gao · 2024
Closest in time.
Beyond probabilities: Unveiling the misalignment in evaluating large language models
C. Lyu, M. Wu, and A. F. Aji · 2024
Closest in time.
MTVQA: Benchmarking multilingual text-centric visual question answering
J. Tang, Q. Liu, Y. Ye, J. Lu, S. Wei, C. Lin, W. Li, M. F. F. B. Mahmood, H. Feng, Z. Zhao, Y. Wang, Y. Liu, H. Liu, X. Bai, and C. Huang · 2024
Closest in time.
Seaeval for multilingual foundation models: From cross-lingual alignment to cultural reasoning, 2024
B. Wang, Z. Liu, X. Huang, F. Jiao, Y. Ding, A. Aw, and N. F. Chen · 2024
Closest in time.