Fetching the paper…
Reading the bibliography…
Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs).
The creative classroom: The role of space and place toward facilitating creativity
Scott A Warner and Kerri L Myers · 2009
Earlier work this paper cites.
Play, imagination, and creativity: A brief literature review
Kuan Chen Tsai · 2012
Earlier work this paper cites.
Lateral thinking in managerial decision making through six thinking hats technique
PS Aithal and PM Kumar · 2017
Earlier work this paper cites.
The effect of problem-based learning on lateral thinking skills
Romy Faisal Mustofa and Yeni Ratna Hidayah · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Naming unrelated words predicts creativity
Jay A Olson, Johnny Nahas, Denis Chmoulevitch, Simon J Cropper, and Margaret E Webb · 2021
Earlier work this paper cites.
Enhancing creativity as innovation via asynchronous crowdwork
Pradeep Kumar Murukannaiah, Nirav Ajmeri, and Munindar P Singh · 2022
Earlier work this paper cites.
Probing the creativity of large language models: Can models produce divergent semantic association?
Honghua Chen and Nai Ding · 2023
Earlier work this paper cites.
Shulin Huang, Shirong Ma, Yinghui Li, Mengzuo Huang, Wuhe Zou, Weidong Zhang, and Hai-Tao Zheng · 2023
Earlier work this paper cites.
Brainteaser: Lateral thinking puzzles for large language models
Yifan Jiang, Filip Ilievski, Kaixin Ma, and Zhivar Sourati · 2023
Earlier work this paper cites.
Brainstorm, then select: a generative language model improves its creativity score
Douglas Summers-Stay, Clare R Voss, and Stephanie M Lukin · 2023
Cited alongside, same era.
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen, Y. Wu, and Zhifang Sui · 2023
Cited alongside, same era.
Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, et al · 2024
Cited alongside, same era.
George Arthur Baker, Ankush Raut, Sagi Shaier, Lawrence E Hunter, and Katharina von der Wense · 2024
Cited alongside, same era.
Mathchat: Benchmarking mathematical reasoning and instruction following in multi-turn interactions
Zhenwen Liang, Dian Yu, Wenhao Yu, Wenlin Yao, Zhihan Zhang, Xiangliang Zhang, and Dong Yu · 2024
Later among the works it cites.
Li-Chun Lu, Shou-Jen Chen, Tsung-Min Pai, Chan-Hung Yu, Hung-yi Lee, and Shao-Hua Sun · 2024
Later among the works it cites.
Guangya Wan, Yuqi Wu, Jie Chen, and Sheng Li · 2024
Later among the works it cites.
Fsm: A finite state machine based zero-shot prompting paradigm for multi-hop question answering
Xiaochen Wang, Junqing He, Yiru Wang, Xiangdi Meng, Kunhao Pan, Zhifang Sui, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antoine Bellemare-Pepin, François Lespinasse, Philipp Thölke, Yann Harel, Kory Mathewson, Jay A. Olson, Yoshua Bengio, and Karim Jerbi · 2024
Cited alongside, same era.
Beyond prompts: Dynamic conversational benchmarking of large language models
David Castillo-Bolado, Joseph Davidson, Finlay Gray, and Marek Rosa · 2024
Cited alongside, same era.
Art or artifice? large language models and the false promise of creativity, 2024
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu · 2024
Cited alongside, same era.
Weak-eval-strong: Evaluating and eliciting lateral thinking of llms with situation puzzles
Qi Chen, Bowen Zhang, Gang Wang, and Qi Wu · 2024
Cited alongside, same era.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Yuanzhuo Wang, and Jian Guo · 2024
Cited alongside, same era.
Ideabench: Benchmarking large language models for research idea generation
Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Eric Xie, Stefan Bekiranov, and Aidong Zhang · 2024
Cited alongside, same era.
Babilong: Testing the limits of llms with long context reasoning-in-a-haystack
Yury Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Cited alongside, same era.
Zhe Xu, Jiasheng Ye, Xiaoran Liu, Xiangyang Liu, Tianxiang Sun, Zhigeng Liu, Qipeng Guo, Linlin Li, Qun Liu, Xuanjing Huang, et al · 2024
Later among the works it cites.
Turtlebench: Evaluating top language models via real-world yes/no puzzles
Qingchen Yu, Shichao Song, Ke Fang, Yunfeng Shi, Zifan Zheng, Hanyu Wang, Simin Niu, and Zhiyu Li · 2024
Later among the works it cites.
Sangwon Yu, Ik-hwan Kim, Jongyoon Song, Saehyung Lee, Junsung Park, and Sungroh Yoon · 2024
Later among the works it cites.
Humanizing llms: A survey of psychological measurements with tools, datasets, and human-agent applications, 2025
Wenhan Dong, Yuemeng Zhao, Zhen Sun, Yule Liu, Zifan Peng, Jingyi Zheng, Zongmin Zhang, Ziyi Zhang, Jun Wu, Ruiming Wang, Shengmin Xu, Xinyi Huang, and Xinlei He · 2025
Closest in time.
Textgames: Learning to self-play text-based puzzle games via language model reasoning
Frederikus Hudi, Genta Indra Winata, Ruochen Zhang, and Alham Fikri Aji · 2025
Closest in time.
Solving situation puzzles with large language model and external reformulation
Kun Li, Xinwei Chen, Tianyou Song, Chengrui Zhou, Zhuoran Liu, Zhenyan Zhang, Jiangjian Guo, and Qing Shan · 2025
Closest in time.
Assessing and understanding creativity in large language models
Yunpu Zhao, Rui Zhang, Wenyi Li, and Ling Li · 2025
Closest in time.