Fetching the paper…
Reading the bibliography…
We study improving social conversational agents by learning from natural dialogue between users and a deployed model, without extra annotations.
Multi-turn beam search for neural dialogue modeling
Ilia Kulikov, Jason Lee, and Kyunghyun Cho. 2019 · 1906
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon. 1992 · 1992
Earlier work this paper cites.
What is user engagement? A conceptual framework for defining user engagement with technology
Heather L O’Brien and Elaine G Toms. 2008 · 2008
Earlier work this paper cites.
A survey on dialogue systems: Recent advances and new frontiers
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017 · 2017
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2018 · 2018
Earlier work this paper cites.
Topic-based evaluation for conversational bots
Fenfei Guo, Angeliki Metallinou, Chandra Khatri, Anirudh Raju, Anu Venkatesh, and Ashwin Ram. 2018 · 2018
Earlier work this paper cites.
Advancing the state of the art in open domain dialog systems through the Alexa prize
Chandra Khatri, Behnam Hedayatnia, Anu Venkatesh, Jeff Nunn, Yi Pan, Qing Liu, Han Song, Anna Gottardi, Sanjeev Kwatra, Sanju Pancholi, et al. 2018 · 2018
Earlier work this paper cites.
Neural response ranking for social conversation: A data-efficient approach
Igor Shalyminov, Ondřej Dušek, and Oliver Lemon. 2018 · 2018
Earlier work this paper cites.
Aiming to know you better perhaps makes me a more engaging dialogue partner
Yury Zemlyanskiy and Fei Sha. 2018 · 2018
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
The second conversational intelligence challenge (ConvAI2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2020 · 2020
Earlier work this paper cites.
Unsupervised evaluation of interactive dialog with DialoGPT
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Cited alongside, same era.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Cited alongside, same era.
Assessing political prudence of open-domain chatbots
Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, and Pascale Fung. 2021 · 2021
Cited alongside, same era.
Understanding and predicting user dissatisfaction in a neural generative chatbot
Abigail See and Christopher Manning. 2021 · 2021
Cited alongside, same era.
Småprat: Dialogpt for natural language generation of swedish dialogue by transfer learning
Oluwatosin Adewumi, Rickard Brännvall, Nosheen Abid, Maryam Pahlavan, Sana Sabah Sabry, Foteini Liwicki, and Marcus Liwicki. 2022 · 2022
Cited alongside, same era.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Later among the works it cites.
The CRINGE loss: Learning what language not to model
Leonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, and Jason Weston. 2023 · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
MERCY: Multiple response ranking concurrently in realistic open-domain conversational systems
Sarik Ghazarian, Behnam Hedayatnia, Di Jin, Sijia Liu, Nanyun Peng, Yang Liu, and Dilek Hakkani-Tur. 2023 · 2023
Closest in time.
ChatGPT outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Cited alongside, same era.
TweetNLP: Cutting-Edge Natural Language Processing for Social Media
Jose Camacho-Collados, Kiamehr Rezaee, Talayeh Riahi, Asahi Ushio, Daniel Loureiro, Dimosthenis Antypas, Joanne Boisson, Luis Espinosa-Anke, Fangyu Liu, Eugenio Martínez-Cámara, et al. 2022 · 2022
Cited alongside, same era.
Model accessible via https://huggingface.co/j-hartmann/emotion-english-distilroberta-base
Jochen Hartmann. 2022 · 2022
Cited alongside, same era.
Da Ju, Jing Xu, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Cited alongside, same era.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Cited alongside, same era.
Amortized noisy channel neural machine translation
Richard Yuanzhe Pang, He He, and Kyunghyun Cho. 2022 · 2022
Cited alongside, same era.
When life gives you lemons, make cherryade: Converting feedback from bad responses into good labels
Weiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, and Jing Xu. 2022 · 2022
Cited alongside, same era.
Robert P. Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Joe Liu, Fritz Cremer, Valentin Assassi, Christie-Carol Beauchamp, Xiaoding Lu, Thomas Rialan, and William Beauchamp. 2023 · 2023
Closest in time.
A framework for vision-language warm-up tasks in multimodal dialogue models
Jaewook Lee, Seongsik Park, Seong-Heum Park, Hongjin Kim, and Harksoo Kim. 2023 · 2023
Closest in time.
Reward gaming in conditional text generation
Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur Parikh, and He He. 2023 · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Closest in time.
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023 · 2023
Closest in time.
Improving open language models by learning from organic interactions
Jing Xu, Da Ju, Joshua Lane, Mojtaba Komeili, Eric Michael Smith, Megan Ung, Morteza Behrooz, William Ngan, Rashel Moritz, Sainbayar Sukhbaatar, et al. 2023 · 2023
Closest in time.
System-level natural language feedback
Weizhe Yuan, Kyunghyun Cho, and Jason Weston. 2023 · 2023
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.