Fetching the paper…
Reading the bibliography…
We explore the creative problem-solving capabilities of modern LLMs in a novel constrained setting.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
On problem-solving
Karl Duncker and Lynne S Lees. 1945 · 1945
Earlier work this paper cites.
Search reduction in hierarchical problem solving
Craig A Knoblock. 1991 · 1991
Earlier work this paper cites.
Social, environmental, and developmental issues and creativity
Beth A Hennessey. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Beyond big and little: The four c model of creativity
James C Kaufman and Ronald A Beghetto. 2009 · 2009
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Phyre: A new benchmark for physical reasoning
Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, and Ross Girshick. 2019 · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych. 2020 · 2020
Earlier work this paper cites.
Prost: Physical reasoning about objects through space and time
Stéphane Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann. 2021 · 2021
Cited alongside, same era.
PTR: A benchmark for part-based conceptual, relational, and physical reasoning
Yining Hong, Li Yi, Joshua B. Tenenbaum, Antonio Torralba, and Chuang Gan. 2021 · 2021
Cited alongside, same era.
Help me write a poem: Instruction tuning as a vehicle for collaborative poetry writing
Tuhin Chakrabarty, Vishakh Padmakumar, and He He. 2022 · 2022
Cited alongside, same era.
Katherine M Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B Tenenbaum. 2022 · 2022
Cited alongside, same era.
AmbiPun: Generating humorous puns with ambiguous context
Anirudh Mittal, Yufei Tian, and Nanyun Peng. 2022 · 2022
Model card and evaluations for claude models
Anthropic. 2023 · 2023
Closest in time.
Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D Hwang, Xiang Lorraine Li, Hirona J Arai, Soumya Sanyal, Keisuke Sakaguchi, Xiang Ren, and Yejin Choi. 2023 · 2023
Closest in time.
Ideas are dimes a dozen: Large language models for idea generation in innovation
Karan Girotra, Lennart Meincke, Christian Terwiesch, and Karl T Ulrich. 2023 · 2023
Closest in time.
Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2023 · 2023
Closest in time.
Best humans still outperform artificial intelligence in a creative divergent thinking task
Mika Koivisto and Simone Grassini. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Introducing chatgpt
OpenAI. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023 · 2023
Cited alongside, same era.
Creativity: Yesterday, today and tomorrow
Joy P Guilford. 1967a
Cited in the paper.
The nature of human intelligence
Joy Paul Guilford. 1967b
Cited in the paper.
We’re afraid language models aren’t modeling ambiguity
Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah Smith, and Yejin Choi. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Unsupervised melody-to-lyrics generation
Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, and Nanyun Peng. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Newton: Are large language models capable of physical reasoning?
Yi Ru Wang, Jiafei Duan, Dieter Fox, and Siddhartha Srinivasa. 2023 · 2023
Closest in time.