Fetching the paper…
Reading the bibliography…
While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data.
Lateral thinking
E. De Bono and E. Zimbalist · 1970
Earlier work this paper cites.
Lateral thinking and technology education
S. Waks · 1997
Earlier work this paper cites.
Eeg complexity and performance measures of creative thinking
M. Mölle, L. Marshall, B. Wolf, H. L. Fehm, and J. Born · 1999
Earlier work this paper cites.
Natural language inference
B. MacCartney · 2009
Earlier work this paper cites.
Development and validation of team creativity measures: A complex systems perspective
H. Jiang and Q.-p. Zhang · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Thinking, fast and slow
K. Daniel · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
R. Speer, J. Chin, and C. Havasi · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph
A. Saha, V. Pahuja, M. Khapra, K. Sankaranarayanan, and S. Chandar · 2018
Earlier work this paper cites.
Comet: Commonsense transformers for automatic knowledge graph construction
A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, and Y. Choi · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
S. Reddy, D. Chen, and C. D. Manning · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Atomic: An atlas of machine commonsense for if-then reasoning
M. Sap, R. Le Bras, E. Allaway, C. Bhagavatula, N. Lourie, H. Rashkin, B. Roof, N. A. Smith, and Y. Choi · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Joint detection and location of english puns
Y. Zou and W. Lu · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Later among the works it cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al · 2023
Later among the works it cites.
S. Huang, S. Ma, Y. Li, M. Huang, W. Zou, W. Zhang, and H.-T. Zheng · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge
B. Y. Lin, Z. Wu, Y. Yang, D.-H. Lee, and X. Ren · 2021
Cited alongside, same era.
Semeval 2021 task 7: Hahackathon, detecting and rating humor and offense
J. Meaney, S. Wilson, L. Chiruzzo, A. Lopez, and W. Magdy · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
Mmdialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation
J. Feng, Q. Sun, C. Xu, P. Zhao, Y. Yang, C. Tao, D. Zhao, and Q. Lin · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Cited alongside, same era.
Later among the works it cites.
Brainteaser: Lateral thinking puzzles for large language model
Y. Jiang, F. Ilievski, and K. Ma · 2023
Later among the works it cites.
Openassistant conversations-democratizing large language model alignment
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. Nguyen, O. Stanley, R. Nagyfi, et al · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Later among the works it cites.
Gaia: a benchmark for general ai assistants
G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom · 2024
Closest in time.
utebc-nlp at semeval-2024 task 9: Can llms be lateral thinkers?
P. Sadeghi, A. Abaskohi, and Y. Yaghoobzadeh · 2024
Closest in time.
Missed connections: Lateral thinking puzzles for large language models
G. Todd, T. Merino, S. Earle, and J. Togelius · 2024
Closest in time.
Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation
S. Zhong, Z. Huang, S. Gao, W. Wen, L. Lin, M. Zitnik, and P. Zhou · 2024
Closest in time.