Fetching the paper…
Reading the bibliography…
How can we train a dialog model to produce better conversations by learning from human feedback, without the risk of humans teaching it harmful chat behaviors? We start by hosting models online, and gather human feedback from real-time, open-ended conversations, which we then use to train and improve the models using offline reinforcement learning (RL).
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare. 2019 · 1901
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019 · 1901
Earlier work this paper cites.
Crossnorm: Normalization for off-policy td reinforcement learning
Aditya Bhatt, Max Argus, Artemij Amiranashvili, and Thomas Brox. 2019 · 1902
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine. 2019 · 1906
Earlier work this paper cites.
Happybot: Generating empathetic dialogue responses by improving user experience look-ahead
Jamin Shin, Peng Xu, Andrea Madotto, and Pascale Fung. 2019 · 1906
Earlier work this paper cites.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. 2019 · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Logic and conversation
Herbert P Grice. 1975 · 1975
Earlier work this paper cites.
Stochastic optimal control
Robert F Stengel. 1986 · 1986
Earlier work this paper cites.
Laughter
Robert R Provine. 1996 · 1996
Earlier work this paper cites.
Functions of humor in the conversations of men and women
Jennifer Hay. 2000 · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup. 2000 · 2000
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade. 2002 · 2002
Earlier work this paper cites.
Where to look: a study of human-robot engagement
Candace L Sidner, Cory D Kidd, Christopher Lee, and Neal Lesh. 2004 · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller. 2005 · 2005
Earlier work this paper cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu. 2020 · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov. 2007 · 2007
Earlier work this paper cites.
Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets
James Henderson, Oliver Lemon, and Kallirroi Georgila. 2008 · 2008
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun. 2010 · 2010
Earlier work this paper cites.
Active listening in peer interviews: The influence of message paraphrasing on perceptions of listening skill
Harry Weger Jr, Gina R Castle, and Melissa C Emmett. 2010 · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurčíček, Blaise Thomson, Kai Yu, and Steve Young. 2011 · 2011
Cited alongside, same era.
Language style matching predicts relationship initiation and stability
Molly E Ireland, Richard B Slatcher, Paul W Eastwick, Lauren E Scissors, Eli J Finkel, and James W Pennebaker. 2011 · 2011
Cited alongside, same era.
Listening competence in initial interactions i: Distinguishing between what listening is and what listeners do
Graham D Bodie, Kellie St. Cyr, Michelle Pence, Michael Rold, and James Honeycutt. 2012 · 2012
Cited alongside, same era.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S Sutton. 2012 · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper. 2012 · 2012
Cited alongside, same era.
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm
Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017 · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Later among the works it cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E Turner, and Douglas Eck. 2017 · 2017
Later among the works it cites.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, and Dan Jurafsky. 2017b · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar. 2012 · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
The role of “active listening” in informal helping conversations: Impact on perceptions of listener helpfulness, sensitivity, and supportiveness and discloser emotional improvement
Graham D Bodie, Andrea J Vickery, Kaitlin Cannava, and Susanne M Jones. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Openai gym
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management
Pei-Hao Su, Paweł Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young. 2017 · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller. 2018 · 2018
Later among the works it cites.
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018 · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018 · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Dialogue generation: From imitation learning to inverse reinforcement learning
Ziming Li, Julia Kiseleva, and Maarten de Rijke. 2018 · 2018
Later among the works it cites.
A hierarchical latent structure for variational conversation modeling
Yookoon Park, Jaemin Cho, and Gunhee Kim. 2018 · 2018
Later among the works it cites.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gokhan Tur. 2018 · 2018
Later among the works it cites.
Sentiment adaptive end-to-end dialog systems
Weiyan Shi and Zhou Yu. 2018 · 2018
Later among the works it cites.
The design and implementation of xiaoice, an empathetic social chatbot
Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2018 · 2018
Later among the works it cites.
Approximating interactive human evaluation with self-play for open-domain dialog systems
Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Agata Lapedriza, and Rosalind Picard. 2019 · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill. 2019 · 2019
Later among the works it cites.
Hierarchical reinforcement learning for open-domain dialog
Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, and Rosalind Picard. 2019 · 2019
Later among the works it cites.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Later among the works it cites.
Unsupervised evaluation of interactive dialog with dialogpt
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Closest in time.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018 · 2069
Closest in time.