Fetching the paper…
Reading the bibliography…
We explore the use of Large Language Model (LLM-based) chatbots to power recommender systems.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
A tutorial on partially observable markov decision processes
Michael L Littman · 2009
Earlier work this paper cites.
Query reformulation using anchor text
Van Dang and Bruce W Croft · 2010
Earlier work this paper cites.
The communicative function of ambiguity in language
Steven T Piantadosi, Harry Tily, and Edward Gibson · 2012
Earlier work this paper cites.
Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects
Jim McCambridge, John Witton, and Diana R Elbourne · 2014
Earlier work this paper cites.
Towards conversational recommender systems
Konstantina Christakopoulou, Filip Radlinski, and Katja Hofmann · 2016
Earlier work this paper cites.
Information gathering actions over human internal state
Dorsa Sadigh, S Shankar Sastry, Sanjit A Seshia, and Anca Dragan · 2016
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
Lerrel Pinto, Marcin Andrychowicz, Peter Welinder, Wojciech Zaremba, and Pieter Abbeel · 2018
Earlier work this paper cites.
Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information
Sudha Rao and Hal Daumé III · 2018
Earlier work this paper cites.
Query expansion techniques for information retrieval: a survey
Hiteshwar Kumar Azad and Akshay Deepak · 2019
Earlier work this paper cites.
Coached conversational preference elicitation: A case study in understanding movie preferences
Filip Radlinski, Krisztian Balog, Bill Byrne, and Karthik Krishnamoorthi · 2019
Earlier work this paper cites.
Analyzing and learning from user interactions for search clarification
Hamed Zamani, Bhaskar Mitra, Everest Chen, Gord Lueck, Fernando Diaz, Paul N Bennett, Nick Craswell, and Susan T Dumais · 2020
Cited alongside, same era.
Ask what’s missing and what’s useful: Improving Clarification Question Generation using Global Knowledge
Bodhisattwa Prasad Majumder, Sudha Rao, Michel Galley, and Julian McAuley · 2021
Cited alongside, same era.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al · 2022
Active prompting with chain-of-thought for large language models
Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang · 2023
Later among the works it cites.
Large language models as zero-shot conversational recommenders
Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley · 2023
Later among the works it cites.
OpenAssistant conversations–democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al · 2023
Later among the works it cites.
Decision-oriented dialogue for human-ai collaboration
Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Hierarchical conversational preference elicitation with bandit feedback
Jinhang Zuo, Songwen Hu, Tong Yu, Shuai Li, Handong Zhao, and Carlee Joe-Wong · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
A categorical archive of chatgpt failures
Ali Borji · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
LLF-Bench: Benchmark for interactive learning from language feedback
Ching-An Cheng, Andrey Kolobov, Dipendra Misra, Allen Nie, and Adith Swaminathan · 2023
Cited alongside, same era.
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Hindsight learning for mdps with exogenous inputs
Sean R Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall, Hugo De Oliveira Barbalho, Jingling Li, Jennifer Neville, Ishai Menache, and Adith Swaminathan · 2023
Later among the works it cites.
A long way to go: Investigating length correlations in RLHF
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Later among the works it cites.
The metacognitive demands and opportunities of generative AI
Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel · 2023
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
MPNet-base-v2, 2024
HuggingFace · 2024
Closest in time.
Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models
Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, and Jaime Teevan · 2024
Closest in time.