Fetching the paper…
Reading the bibliography…
Maintaining a consistent persona is a key quality for any open domain dialogue system.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder Singh. 2000 · 2000
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, G. Tucker, and Justin Fu. 2020 · 2005
Earlier work this paper cites.
Predicting pragmatic reasoning in language games
Michael C. Frank and Noah D. Goodman. 2012 · 2012
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare. 2016 · 2016
Earlier work this paper cites.
ParlAI: A dialog research software platform
Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017 · 2017
Earlier work this paper cites.
A deep reinforcement learning chatbot
Iulian V. Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian, Taesup Kim, Michael Pieper, Sarath Chandar, Nan Rosemary Ke, Sai Rajeshwar, Alexandre de Brebisson, Jose M. R. Sotelo, Dendi Suhubdy, Vincent Michalski, Alexandre Nguyen, Joelle Pineau, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Importance of search and evaluation strategies in neural dialogue modeling
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum. 2019 · 2019
Cited alongside, same era.
DeepCopy: Grounded response generation with hierarchical pointer networks
Semih Yavuz, Abhinav Rastogi, Guan-Lin Chao, and Dilek Hakkani-Tur. 2019 · 2019
Cited alongside, same era.
Generate, delete and rewrite: A three-stage framework for improving persona consistency of dialogue generation
Haoyu Song, Yan Wang, Wei-Nan Zhang, Xiaojiang Liu, and Ting Liu. 2020 · 2020
Later among the works it cites.
Text generation by learning from demonstrations
Richard Yuanzhe Pang and He He. 2021 · 2021
Later among the works it cites.
Refine and imitate: Reducing repetition and inconsistency in persuasion dialogues via reinforcement learning and human demonstration
Weiyan Shi, Yu Li, Saurav Sahay, and Zhou Yu. 2021 · 2021
Later among the works it cites.
Gpt-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim. 2022 · 2022
Later among the works it cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-centric dialog training via offline reinforcement learning
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2020 · 2020
Cited alongside, same era.
Will I sound like me? improving persona consistency in dialogues through pragmatic self-consciousness
Hyunwoo Kim, Byeongchang Kim, and Gunhee Kim. 2020 · 2020
Cited alongside, same era.
Don’t say that! making inconsistent dialogue unlikely with unlikelihood training
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Cited alongside, same era.
You impress me: Dialogue generation via mutual persona perception
Qian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou, Zixuan Chen, Bin Zhou, and Dongmei Zhang. 2020 · 2020
Cited alongside, same era.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M. Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2020
Cited alongside, same era.
Exploiting persona information for diverse generation of conversational responses
Haoyu Song, Wei-Nan Zhang, Yiming Cui, Dong Wang, and Ting Liu. 2019a
Cited in the paper.
Generating persona consistent dialogues by exploiting natural language inference
Haoyu Song, Weinan Zhang, Jingwen Hu, and Ting Liu. 2019b
Cited in the paper.
Later among the works it cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Y-Lan Boureau, Melanie Kambadur, and Jason Weston. 2022 · 2022
Later among the works it cites.
Context-aware language modeling for goal-oriented dialogue systems
Charlie Snell, Sherry Yang, Justin Fu, Yi Su, and Sergey Levine. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022 · 2022
Later among the works it cites.
CHAI: A CHatbot AI for task-oriented dialogue with offline reinforcement learning
Siddharth Verma, Justin Fu, Sherry Yang, and Sergey Levine. 2022 · 2022
Later among the works it cites.
Offline RL for natural language generation with implicit language q learning
Charlie Victor Snell, Ilya Kostrikov, Yi Su, Sherry Yang, and Sergey Levine. 2023 · 2023
Closest in time.