Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al.
Humor and Laughter: An Anthropological Approach
Surender Apte. 1988 · 1988
Earlier work this paper cites.
Computational humor: An implemented model of puns
Kim Binsted et al. 2006 · 2006
Earlier work this paper cites.
The New Yorker Cartoon Caption Contest , first edition edition
Robert Mankoff. 2008 · 2008
Earlier work this paper cites.
Random walk factoid annotation for collective discourse
Ben King, Rahul Jha, Dragomir Radev, and Robert Mankoff. 2013 · 2013
Earlier work this paper cites.
Next: A system for real-world development, evaluation, and application of active learning
Kevin G Jamieson, Lalit Jain, Chris Fernandez, Nicholas J Glattard, and Rob Nowak. 2015 · 2015
Earlier work this paper cites.
Inside jokes: Identifying humorous cartoon captions
Dafna Shahaf, Eric Horvitz, and Robert Mankoff. 2015 · 2015
Earlier work this paper cites.
Humor in collective discourse: Unsupervised funniness detection in the New Yorker cartoon caption contest
Dragomir Radev, Amanda Stent, Joel Tetreault, Aasish Pappu, Aikaterini Iliakopoulou, Agustin Chanfreau, Paloma de Juan, Jordi Vallmitjana, Alejandro Jaimes, Rahul Jha, and Robert Mankoff. 2016 · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, et al. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022 · 2022
Earlier work this paper cites.
Simulating opinion dynamics with networks of llm-based agents
Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. 2023 · 2023
Cited alongside, same era.
Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest
Jack Hessel, Ana Marasović, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2023 · 2023
Cited alongside, same era.
Chatgpt is fun, but it is not funny! humor is still challenging large language models
Sophie Jentzsch and Kristian Kersting. 2023 · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior. arxiv
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Cited alongside, same era.
Gemini 2.0 flash thinkingg
Google. 2024 · 2024
Later among the works it cites.
Group relative policy optimization for aligning large language models
Daya Guo, Qihao Zhu, Dejian Yang, and Junxiao Song. 2024 · 2024
Later among the works it cites.
Getting serious about humor: Crafting humor datasets with unfunny large language models
Zachary Horvitz et al. 2024 · 2024
Later among the works it cites.
A robot walks into a bar: Can language models serve as creativity support tools for comedy?
Piotr W. Mirowski et al. 2024 · 2024
Later among the works it cites.
Is ai funnier than humans? this study says so — but you be the judge
New York Post. 2024 · 2024
Later among the works it cites.
The last laugh: Exploring the role of humor as a benchmark for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, and Pan Zhou. 2023 · 2023
Cited alongside, same era.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
From persona to personalization: A survey on role-playing language agents
Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, et al. 2024 · 2024
Cited alongside, same era.
The wisdom of partisan crowds: Comparing collective intelligence in humans and llm-based agents
Yun-Shiuan Chuang, Nikunj Harlalka, Siddharth Suresh, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. 2024 · 2024
Cited alongside, same era.
Deepseek
DeepSeek. 2024 · 2024
Cited alongside, same era.
Hello gpt-4o
OpenAI. 2024a
Cited in the paper.
Introducing openai o1-preview
OpenAI. 2024b
Cited in the paper.
Greg Robison. 2024 · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon captioning
Jifan Zhang, Lalit Jain, Yang Guo, Jiayi Chen, Kuan Lok Zhou, Siddharth Suresh, Andrew Wagenmaker, Scott Sievert, Timothy Rogers, Kevin Jamieson, et al. 2024 · 2024
Later among the works it cites.
Text is not all you need: Multimodal prompting helps llms understand humor
Ashwin Baluja. 2025 · 2025
Closest in time.
Openai o3-mini
OpenAI. 2025 · 2025
Closest in time.