Fetching the paper…
Reading the bibliography…
Language models have demonstrated remarkable capabilities on standard benchmarks, yet they struggle increasingly from mode collapse, the inability to generate diverse and novel outputs.
Bleu: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Rank-biased precision for measurement of retrieval effectiveness
Alistair Moffat and Justin Zobel · 2008
Earlier work this paper cites.
A Diversity-Promoting Objective Function for Neural Conversation Models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2016
Earlier work this paper cites.
Deal or No Deal? End-to-End Learning of Negotiation Dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Hierarchical Neural Story Generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Texygen: A Benchmarking Platform for Text Generation Models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Earlier work this paper cites.
Boosting Dialog Response Generation
Wenchao Du and Alan W Black · 2019
Earlier work this paper cites.
Generating Diverse Translations with Sentence Codes
Raphael Shu, Hideki Nakayama, and Kyunghyun Cho · 2019
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2019
Earlier work this paper cites.
Evaluating the state-of-the-art of End-to-End Natural Language Generation: The E2E NLG challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser · 2020
Earlier work this paper cites.
Simultaneous translation and paraphrase for language education
Stephen Mayhew, Klinton Bicknell, Chris Brust, Bill McDowell, Will Monroe, and Burr Settles · 2020
Earlier work this paper cites.
DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Earlier work this paper cites.
Evaluating the Evaluation of Diversity in Natural Language Generation
Guy Tevet and Jonathan Berant · 2021
Earlier work this paper cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al · 2023
Earlier work this paper cites.
Large Language Model (LLM) Bias Index - LLMBI
Abiodun Finbarrs Oketunji, Muhammad Anas, and Deepthi Saina · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Aligning with whom? large language models have gender and racial biases in subjective nlp tasks
Huaman Sun, Jiaxin Pei, Minje Choi, and David Jurgens · 2023
Cited alongside, same era.
WildChat: 1M ChatGPT Interaction Logs in the Wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Detecting mode collapse in language models via narration
Sil Hamilton · 2024
Later among the works it cites.
RewardBench: Evaluating Reward Models for Language Modeling, June 2024
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, L. J. Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
Ziniu Li, Congliang Chen, Tian Xu, Zeyu Qin, Jiancong Xiao, Zhi-Quan Luo, and Ruoyu Sun · 2024
Later among the works it cites.
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs, October 2024
Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, and Yahui Zhou · 2024
Later among the works it cites.
The Llama 3 Herd of Models, July 2024
Llama Team et al · 2024
Later among the works it cites.
GPT-4o System Card
OpenAI · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ai suggestions homogenize writing toward western styles and diminish cultural nuances
Dhruv Agarwal, Mor Naaman, and Aditya Vashistha · 2024
Cited alongside, same era.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic · 2024
Cited alongside, same era.
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
Qi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O’Gorman, Nalini Singh, Andrew McCallum, and Xiang Li · 2024
Cited alongside, same era.
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations, November 2024
Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti · 2024
Cited alongside, same era.
Command R and Command R+ Model Card
Cohere · 2024
Cited alongside, same era.
Angry men, sad women: Large language models reflect gendered stereotypes in emotion attribution
Flor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, and Dirk Hovy · 2024
Cited alongside, same era.
Gemma 2: Improving Open Language Models at a Practical Size, October 2024
Gemma Team et al · 2024
Cited alongside, same era.
Later among the works it cites.
Attributing mode collapse in the fine-tuning of large language models
Laura O’Mahony, Leo Grinsztajn, Hailey Schoelkopf, and Stella Biderman · 2024
Later among the works it cites.
Is temperature the creativity parameter of large language models?
Max Peeperkorn, Tom Kouwenhoven, Daniel G. Brown, and Anna K. Jordanous · 2024
Later among the works it cites.
Growing a tail: Increasing output diversity in large language models
Michal Shur-Ofry, Bar Horowitz-Amsalem, Adir Rahamim, and Yonatan Belinkov · 2024
Later among the works it cites.
The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism, 2024
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin · 2024
Later among the works it cites.
Position: A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi · 2024
Later among the works it cites.
HelpSteer2-Preference: Complementing Ratings with Preferences
Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, and Yi Dong · 2024
Later among the works it cites.
Diverse Preference Optimization, February 2025
Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov · 2025
Closest in time.
2 OLMo 2 Furious, January 2025
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Michal Guerquin, Hamish Ivison, Pang Wei Koh, Jiacheng Liu, Saumya Malik, William Merrill, Lester James V. Miranda, Jacob Morrison, Tyler Murray, Crystal Nam, Valentina Pyatkin, Aman Rangapur, Michael Schmitz, Sam Skjonsberg, David Wadden, Christopher Wilhelm, Michael Wilson, Luke Zettlemoyer, Ali Farhadi, Noah A. Smith, and Hannaneh Hajishirzi · 2025
Closest in time.