Fetching the paper…
Reading the bibliography…
Human-LLM conversations are increasingly becoming more pervasive in peoples' professional and personal lives, yet many users still struggle to elicit helpful responses from LLM Chatbots.
Offline reinforcement learning from human feedback in real-world sequence-to-sequence tasks
Julia Kreutzer, Stefan Riezler, and Carolin Lawrence. 2021 · 2021
Earlier work this paper cites.
Learning a voice-based conversational recommender using offline policy optimization
Francois Mairesse, Zhonghao Luo, and Tao Ye. 2021 · 2021
Earlier work this paper cites.
Characterizing stage-aware writing assistance for collaborative document authoring
Bahareh Sarrafzadeh, Sujay Kumar Jauhar, Michael Gamon, Edward Lank, and Ryen W. White. 2021 · 2021
Earlier work this paper cites.
Elaborative simplification: Content addition and explanation generation in text simplification
Neha Srikanth and Junyi Jessy Li. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023 · 2023
Earlier work this paper cites.
Natural language decompositions of implicit content enable better text representations
Alexander Hoyle, Rupak Sarkar, Pranav Goel, and Philip Resnik. 2023 · 2023
Earlier work this paper cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023 · 2023
Earlier work this paper cites.
Query rewriting in retrieval-augmented large language models
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023 · 2023
Earlier work this paper cites.
FollowupQG: Towards information-seeking follow-up question generation
Yan Meng, Liangming Pan, Yixin Cao, and Min-Yen Kan. 2023 · 2023
Earlier work this paper cites.
Overview of robust and multilingual automatic evaluation metricsfor open-domain dialogue systems at DSTC 11 track 4
Mario Rodríguez-Cantelar, Chen Zhang, Chengguang Tang, Ke Shi, Sarik Ghazarian, João Sedoc, Luis Fernando D’Haro, and Alexander I. Rudnicky. 2023 · 2023
Earlier work this paper cites.
Why johnny can’t prompt: How non-ai experts try (and fail) to design llm prompts
J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023 · 2023
Earlier work this paper cites.
Llmeval: A preliminary study on how to evaluate large language models
Yue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023 · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
StudentEval: A benchmark of student-written prompts for large language models of code
Hannah McLean Babe, Sydney Nguyen, Yangtian Zi, Arjun Guha, Molly Q Feldman, and Carolyn Jane Anderson. 2024 · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, and Anirudh Goyal et. al. 2024 · 2024
Cited alongside, same era.
Why and when llm-based assistants can go wrong: Investigating the effectiveness of prompt-based interactions for software help-seeking
Anjali Khurana, Hariharan Subramonyam, and Parmit K Chilana. 2024 · 2024
Cited alongside, same era.
Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen, and Paul Denny. 2024 · 2024
Cited alongside, same era.
Learning to rewrite prompts for personalized text generation
Cheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong, and Michael Bendersky. 2024a · 2024
Cited alongside, same era.
Leveraging large language models for NLG evaluation: Advances and challenges
Wildfeedback: Aligning llms with in-situ user interactions and feedback
Taiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin, Zexue He, Mengting Wan, Pei Zhou, Sujay Jauhar, Xiaofeng Xu, Xia Song, and Jennifer Neville. 2024 · 2024
Later among the works it cites.
Are large language models capable of generating human-level narratives?
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, and Nanyun Peng. 2024 · 2024
Later among the works it cites.
Tnt-llm: Text mining at scale with large language models
Mengting Wan, Tara Safavi, Sujay Kumar Jauhar, Yujin Kim, Scott Counts, Jennifer Neville, Siddharth Suri, Chirag Shah, Ryen W White, Longqi Yang, et al. 2024 · 2024
Later among the works it cites.
Understanding user experience in large language model interactions
Jiayin Wang, Weizhi Ma, Peijie Sun, Min Zhang, and Jian-Yun Nie. 2024 · 2024
Later among the works it cites.
Hallucination is inevitable: An innate limitation of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, and Shuai Ma. 2024c · 2024
Cited alongside, same era.
Interpretable user satisfaction estimation for conversational systems with large language models
Ying-Chun Lin, Jennifer Neville, Jack W Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, et al. 2024 · 2024
Cited alongside, same era.
Na Liu, Liangyu Chen, Xiaoyu Tian, Wei Zou, Kaijiang Chen, and Ming Cui. 2024 · 2024
Cited alongside, same era.
Contextualized evaluations: Taking the guesswork out of language model evaluations
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth, Mohit Iyyer, Mark Yatskar, and Kyle Lo. 2024 · 2024
Cited alongside, same era.
On the benchmarking of LLMs for open-domain dialogue evaluation
John Mendonça, Alon Lavie, and Isabel Trancoso. 2024a · 2024
Cited alongside, same era.
Soda-eval: Open-domain dialogue evaluation in the age of LLMs
John Mendonça, Isabel Trancoso, and Alon Lavie. 2024b · 2024
Cited alongside, same era.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024 · 2024
Cited alongside, same era.
Llm targeted underperformance disproportionately impacts vulnerable users
Elinor Poole-Dayan, Deb Roy, and Jad Kabbara. 2024 · 2024
Cited alongside, same era.
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Later among the works it cites.
InterrogateLLM: Zero-resource hallucination detection in LLM-generated answers
Yakir Yehuda, Itzik Malkiel, Oren Barkan, Jonathan Weill, Royi Ronen, and Noam Koenigstein. 2024 · 2024
Later among the works it cites.
Wildchat: 1m chatgpt interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 · 2024
Later among the works it cites.
Beyond the safety bundle: Auditing the helpful and harmless dataset
Khaoula Chehbouni, Jonathan Colaço Carr, Yash More, Jackie CK Cheung, and Golnoosh Farnadi. 2025 · 2025
Closest in time.
DeepSeek-AI. 2025 · 2025
Closest in time.
The BiGGen bench: A principled benchmark for fine-grained evaluation of language models with language models
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2025 · 2025
Closest in time.
Openai o3-mini system card
OpenAI. 2025 · 2025
Closest in time.
Adapt: Actively discovering and adapting to preferences for any task
Maithili Patel, Xavier Puig, Ruta Desai, Roozbeh Mottaghi, Sonia Chernova, Joanne Truong, and Akshara Rai. 2025 · 2025
Closest in time.
Judging the judges: A systematic study of position bias in llm-as-a-judge
Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2025 · 2025
Closest in time.
Aligning LLMs with individual preferences via interaction
Shujin Wu, Yi R. Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2025 · 2025
Closest in time.
Language model council: Democratically benchmarking foundation models on highly subjective tasks
Justin Zhao, Flor Miriam Plaza-del Arco, and Amanda Cercas Curry. 2025 · 2025
Closest in time.