Fetching the paper…
Reading the bibliography…
Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts.
The triarchic theory of intelligence
Robert J Sternberg · 1997
Earlier work this paper cites.
Fifty years of creativity research
Richard E Mayer · 1999
Earlier work this paper cites.
Possible brain mechanisms of creativity
Kenneth M Heilman · 2016
Earlier work this paper cites.
Creativity in context: Update to the social psychology of creativity
Teresa M Amabile · 2018
Earlier work this paper cites.
Subcortical structures and visual divergent thinking: a resting-state functional mri analysis
Zhenni Gao, Xiaojin Liu, Delong Zhang, Ming Liu, and Ning Hao · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Putting gpt-3’s creativity to the (alternative uses) test
Claire Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han van der Maas · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2023
Earlier work this paper cites.
Mllm-bench: evaluating multimodal llms with per-sample criteria
Wentao Ge, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, et al · 2023
Earlier work this paper cites.
The originality of machines: Ai takes the torrance test
Erik E Guzik, Christian Byrge, and Christian Gilde · 2023
Earlier work this paper cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2023
Cited alongside, same era.
Are we on the right way for evaluating large vision-language models?
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al · 2024
Cited alongside, same era.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, et al · 2024
Cited alongside, same era.
A confederacy of models: A comprehensive evaluation of llms on creative writing
Paul Williams and Carlos Gómez-Rodríguez · 2024
Later among the works it cites.
Alignmmbench: Evaluating chinese multimodal alignment in large vision-language models
Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, and Yuxiao Dong · 2024
Later among the works it cites.
Llm-evolve: Evaluation for llm’s evolving capability on benchmarks
Jiaxuan You, Mingjie Liu, Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Eric Xie, Stefan Bekiranov, and Aidong Zhang · 2024
Cited alongside, same era.
Simulbench: Evaluating language models with creative simulation tasks
Qi Jia, Xiang Yue, Tianyu Zheng, Jie Huang, and Bill Yuchen Lin · 2024
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al · 2024
Cited alongside, same era.
Do llms agree on the creativity evaluation of alternative uses?
Abdullah Al Rabeyah, Fabrício Góes, Marco Volpe, and Talles Medeiros · 2024
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2024
Cited alongside, same era.
Liveideabench: Evaluating llms’ scientific creativity and idea generation with minimal context
Kai Ruan, Xuan Wang, Jixiang Hong, and Hao Sun · 2024
Cited alongside, same era.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al · 2024
Cited alongside, same era.
Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, et al · 2024
Later among the works it cites.
Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng · 2025
Closest in time.
Llm creative story-writing benchmark
Lech Mazur · 2025
Closest in time.
Prism: A framework for decoupling and assessing the capabilities of vlms
Yuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang, Lin Chen, Songyang Zhang, Jiaqi Wang, Dahua Lin, and Kai Chen · 2025
Closest in time.
Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, et al · 2025
Closest in time.
Rpgbench: Evaluating large language models as role-playing game engines
Pengfei Yu, Dongming Shen, Silin Meng, Jaewon Lee, Weisu Yin, Andrea Yaoyun Cui, Zhenlin Xu, Yi Zhu, Xingjian Shi, Mu Li, et al · 2025
Closest in time.
Redundancy principles for mllms benchmarks, 2025
Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, and Guangtao Zhai · 2025
Closest in time.