Fetching the paper…
Reading the bibliography…
Evaluating the creativity of large language models (LLMs) in story writing is difficult because LLM-generated stories could seemingly look creative but be very similar to some existing stories in their huge and proprietary training corpus.
Encoding specificity and retrieval processes in episodic memory
Endel Tulving and Donald M Thomson. 1973 · 1973
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Earlier work this paper cites.
Creative writing with an ai-powered writing assistant: Perspectives from professional writers
Daphne Ippolito, Ann Yuan, Andy Coenen, and Sehmon Burnam. 2022 · 2022
Earlier work this paper cites.
Augmenting scientific creativity with an analogical search engine
Hyeonsu B Kang, Xin Qian, Tom Hope, Dafna Shahaf, Joel Chan, and Aniket Kittur. 2022 · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
The crowdless future? generative ai and creative problem solving
Leonard Boussioux, Jacqueline Ng Lane, Miaomiao Zhang, Vladimir Jacimovic, and Karim R Lakhani. 2023 · 2023
Earlier work this paper cites.
Speak, memory: An archaeology of books known to chatgpt/gpt-4
Kent K Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Earlier work this paper cites.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Fluid transformers and creative analogies: Exploring large language models’ capacity for augmenting cross-domain analogical creativity
Zijian Ding, Arvind Srinivasan, Stephen MacNeil, and Joel Chan. 2023 · 2023
Earlier work this paper cites.
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Earlier work this paper cites.
A confederacy of models: A comprehensive evaluation of llms on creative writing
Carlos Gómez-Rodríguez and Paul Williams. 2023 · 2023
Earlier work this paper cites.
Sok: Memorization in general-purpose large language models
Valentin Hartmann, Anshuman Suri, Vincent Bindschaedler, David Evans, Shruti Tople, and Robert West. 2023 · 2023
Earlier work this paper cites.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramer, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher Choquette Choo, and Nicholas Carlini. 2023 · 2023
Earlier work this paper cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Earlier work this paper cites.
A long way to go: Investigating length correlations in rlhf
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Cited alongside, same era.
The generative ai paradox:“what it can create, it may not understand
Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, et al. 2023 · 2023
Cited alongside, same era.
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim. 2023 · 2023
Cited alongside, same era.
Exploring precision and recall to assess the quality and diversity of llms
Florian Le Bronnec, Alexandre Vérine, Benjamin Negrevergne, Yann Chevaleyre, and Alexandre Allauzen. 2024 · 2024
Closest in time.
How ai processing delays foster creativity: Exploring research question co-creation with an llm-based agent
Yiren Liu, Si Chen, Haocong Cheng, Mengxia Yu, Xiao Ran, Andrew Mo, Yiliu Tang, and Yun Huang. 2024 · 2024
Closest in time.
Llm comparative assessment: Zero-shot nlg evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Closest in time.
Benchmarking language model creativity: A case study on code generation
Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, and Daniel Khashabi. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shunyu Yao, Howard Chen, Austin W Hanjie, Runzhe Yang, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023 · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 · 2023
Cited alongside, same era.
Evaluating creativity support tools via homogenization analysis
Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. 2024 · 2024
Cited alongside, same era.
Divergent creativity in humans and large language models
Antoine Bellemare-Pepin, François Lespinasse, Philipp Thölke, Yann Harel, Kory Mathewson, Jay A Olson, Yoshua Bengio, and Karim Jerbi. 2024 · 2024
Cited alongside, same era.
Art or artifice? large language models and the false promise of creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024 · 2024
Cited alongside, same era.
Tong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, and Pang Wei Koh. 2024 · 2024
Cited alongside, same era.
Instructeval: Towards holistic evaluation of instruction-tuned large language models
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2024 · 2024
Cited alongside, same era.
Guillermo Marco, Julio Gonzalo, Ramón del Castillo, and María Teresa Mateo Girona. 2024 · 2024
Closest in time.
A robot walks into a bar: Can language models serve as creativity supporttools for comedy? an evaluation of llms’ humour alignment with comedians
Piotr Mirowski, Juliette Love, Kory Mathewson, and Shakir Mohamed. 2024 · 2024
Closest in time.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev. 2024 · 2024
Closest in time.
Suri: Multi-constraint instruction following for long-form text generation
Chau Minh Pham, Simeng Sun, and Mohit Iyyer. 2024 · 2024
Closest in time.
Infobench: Evaluating instruction following ability in large language models
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
Are large language models capable of generating human-level narratives?
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, and Nanyun Peng. 2024 · 2024
Closest in time.
Kaiwen Wei, Jingyuan Zhang, Hongzhi Zhang, Fuzheng Zhang, Di Zhang, Li Jin, and Yue Yu. 2024 · 2024
Closest in time.
Benchmarking complex instruction-following with multiple constraints composition
Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu, Hao Huang, Jinfeng Zhou, Wenchuang Li, Binxin Hu, Wendy Gao, Jiaxin Xu, et al. 2024 · 2024
Closest in time.
Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J Su. 2024 · 2024
Closest in time.
Cfbench: A comprehensive constraints-following benchmark for llms
Tao Zhang, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, et al. 2024 · 2024
Closest in time.
Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, and Pan Zhou. 2024 · 2024
Closest in time.