Fetching the paper…
Reading the bibliography…
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Evaluating hallucinations in chinese large language models
Q. Cheng, T. Sun, W. Zhang, S. Wang, X. Liu, M. Zhang, J. He, M. Huang, Z. Yin, K. Chen, et al · 2023
Earlier work this paper cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He · 2023
Earlier work this paper cites.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2023
Earlier work this paper cites.
Gpt-4 technical report
OpenAI · 2023
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2023
Earlier work this paper cites.
Baichuan 2: Open large-scale language models
A. Yang, B. Xiao, B. Wang, B. Zhang, C. Bian, C. Yin, C. Lv, D. Pan, D. Wang, D. Yan, et al · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Cited alongside, same era.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2024
Later among the works it cites.
Chinese safetyqa: A safety short-form factuality benchmark for large language models, 2024
Y. Tan, B. Zheng, B. Zheng, K. Cao, H. Jing, J. Wei, J. Liu, Y. He, W. Su, X. Zhu, and B. Zheng · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models
S. Tonmoy, S. Zaman, V. Jain, A. Rani, V. Rawte, A. Chadha, and A. Das · 2024
Later among the works it cites.
Measuring short-form factuality in large language models. 2024
J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus · 2024
Later among the works it cites.
Measuring short-form factuality in large language models
J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Are we on the right way for evaluating large vision-language models?
L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al · 2024
Cited alongside, same era.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
Seed-bench: Benchmarking multimodal large language models
B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan · 2024
Cited alongside, same era.
Xformparser: A simple and effective multimodal multilingual semi-structured form parser
X. Cheng, H. Zhang, J. Yang, X. Li, W. Zhou, K. Wu, F. Liu, W. Zhang, T. Sun, T. Li, et al
Cited in the paper.
Sviptr: Fast and efficient scene text recognition with vision permutable extractor
X. Cheng, W. Zhou, X. Li, J. Yang, H. Zhang, T. Sun, W. Zhang, Y. Mai, T. Li, X. Chen, et al
Cited in the paper.
Chinese simpleqa: A chinese factuality evaluation for large language models, 2024a
Y. He, S. Li, J. Liu, Y. Tan, W. Wang, H. Huang, X. Bu, H. Guo, C. Hu, B. Zheng, Z. Lin, X. Liu, D. Sun, S. Lin, Z. Zheng, X. Zhu, W. Su, and B. Zheng
Cited in the paper.
Chinese simpleqa: A chinese factuality evaluation for large language models
Y. He, S. Li, J. Liu, Y. Tan, W. Wang, H. Huang, X. Bu, H. Guo, C. Hu, B. Zheng, et al
Cited in the paper.
Later among the works it cites.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
G. Zhang, X. Du, B. Chen, Y. Liang, T. Luo, T. Zheng, K. Zhu, Y. Cheng, C. Xu, S. Guo, H. Zhang, X. Qu, J. Wang, R. Yuan, Y. Li, Z. Wang, Y. Liu, Y.-H. Tsai, F. Zhang, C. Lin, W. Huang, W. Chen, and J. Fu · 2024
Later among the works it cites.
C. Zou, X. Guo, R. Yang, J. Zhang, B. Hu, and H. Zhang · 2024
Later among the works it cites.
Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding
H. Li, J. Chen, Z. Wei, S. Huang, T. Hui, J. Gao, X. Wei, and S. Liu · 2025
Closest in time.