Fetching the paper…
Reading the bibliography…
Evaluating creative writing generated by large language models (LLMs) remains challenging because open-ended narratives lack ground truths.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
The cults of “research” and “creativity”
Jacques Barzun · 1960
Earlier work this paper cites.
An analysis of creativity
Mel Rhodes · 1961
Earlier work this paper cites.
Social psychology of creativity: A consensual assessment technique
Teresa M. Amabile · 1982
Earlier work this paper cites.
Assessing Writing
Sara Cushing Weigle · 2002
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2009
Earlier work this paper cites.
Linguistic features of writing quality
Danielle S McNamara, Scott A Crossley, and Philip M McCarthy · 2010
Earlier work this paper cites.
Factors Affecting Upvoting Intention on Social Bookmarking Sites
Kasra Kassaeyan · 2016
Earlier work this paper cites.
Jointly measuring diversity and quality in text generation models
Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah · 2019
Earlier work this paper cites.
Beyond subjective judgments: Predicting evaluations of creative writing from computational linguistic features
Claire M Zedelius, Caitlin Mills, and Jonathan W Schooler · 2019
Earlier work this paper cites.
On the relation between quality–diversity evaluation and distribution-fitting goal in text generation
Jianing Li, Yanyan Lan, Jiafeng Guo, and Xueqi Cheng · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
The Practice of Creative Writing: A Guide for Students
Heather Sellers · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Cited alongside, same era.
How interesting and coherent are the stories generated by a large-scale neural language model? comparing human and automatic evaluations of machine-generated text
Dominic Callan and Jennifer Foster · 2023
Cited alongside, same era.
Poetry will not optimize; or, what is literature to ai?
Michele Elam · 2023
Cited alongside, same era.
Stanford human preferences dataset, 2023
Training software engineering agents and verifiers with swe-gym
Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang · 2024
Later among the works it cites.
Llm-as-a-judge and reward model: What they can and cannot do, 2024
Guijin Son, Hyunwoo Ko, Hoyoung Lee, Yewon Kim, and Seunghyeok Hong · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V Chawla, and Xiangliang Zhang · 2024
Later among the works it cites.
Star: Self-taught reasoner bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D Goodman · 2024
Later among the works it cites.
Generative verifiers: Reward modeling as next-token prediction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kawin Ethayarajh, Heidi Zhang, Yizhong Wang, and Dan Jurafsky · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Jonas Schneider, Jonathan Raiman, Ian Tenney, Nitish Gupta, Diya Raghu, Douwe Kiela, and Lazaros Polymenakos · 2023
Cited alongside, same era.
Comparison of evaluation metrics for short story generation
Ponrudee Netisopakul and Usanisa Taoto · 2023
Cited alongside, same era.
Large language models are not fair evaluators, 2023
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng et al · 2023
Cited alongside, same era.
Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text, 2024
Sher Badshah and Hassan Sajjad · 2024
Cited alongside, same era.
Art or artifice? large language models and the false promise of creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu · 2024
Cited alongside, same era.
Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal · 2024
Later among the works it cites.
The user demographics of reddit: The official app
Abhinav Agrawal · 2025
Closest in time.
Modifying large language model post-training for diverse creative writing, 2025
John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yuqian Sun, and Max Kreminski · 2025
Closest in time.
Think, prune, train, improve: Scaling reasoning without scaling models, 2025
Caia Costello, Simon Guo, Anna Goldie, and Azalia Mirhoseini · 2025
Closest in time.
Reddit user age, gender, & demographics (2025)
Fabio Duarte · 2025
Closest in time.
Style outweighs substance: Failure modes of llm judges in alignment benchmarking, 2025
Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, and John P. Dickerson · 2025
Closest in time.
Automated creativity evaluation for large language models: A reference-based approach, 2025
Ruizhe Li, Chiwei Zhu, Benfeng Xu, Xiaorui Wang, and Zhendong Mao · 2025
Closest in time.
Hui Wei, Shenghua He, Tian Xia, Fei Liu, Andy Wong, Jingyang Lin, and Mei Han · 2025
Closest in time.
Optimizing generative ai by backpropagating language model feedback
Mert Yuksekgonul et al · 2025
Closest in time.
The Education of the Creative Writing Teacher: A Study of Conceptions of Creative Writing Pedagogy in Higher Education
Rebecca Manery · 2027
Closest in time.