Fetching the paper…
Reading the bibliography…
Most existing approaches for goal-oriented dialogue policy learning used reinforcement learning, which focuses on the target agent policy and simply treat the opposite agent policy as part of the environment.
Budgeted policy learning for task-oriented dialogue systems
Zhirui Zhang, Xiujun Li, Jianfeng Gao, and Enhong Chen. 2019b · 1906
Earlier work this paper cites.
Guided dialog policy learning: Reward estimation for multi-domain task-oriented dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 1908
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff. 1978 · 1978
Earlier work this paper cites.
Folk psychology as simulation
Robert M Gordon. 1986 · 1986
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997 · 1997
Earlier work this paper cites.
Mirror neurons and the simulation theory of mind-reading
Vittorio Gallese and Alvin Goldman. 1998 · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies
Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. 2006 · 2006
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Partially observable markov decision processes for spoken dialog systems
Jason D Williams and Steve Young. 2007 · 2007
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Earlier work this paper cites.
Learning from 26 languages: Program management and science in the babel program
Mary Harper. 2014 · 2014
Earlier work this paper cites.
Policy learning for domain selection in an extensible multi-domain spoken dialogue system
Zhuoran Wang, Hongliang Chen, Guanchun Wang, Hao Tian, Hua Wu, and Haifeng Wang. 2014 · 2014
Earlier work this paper cites.
Policy committee for adaptation in multi-domain spoken dialogue systems
M Gašić, N Mrkšić, Pei-hao Su, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Cited alongside, same era.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Layla El Asri, Jing He, and Kaheer Suleman. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning for multi-domain dialogue systems
Heriberto Cuayáhuitl, Seunghak Yu, Ashley Williamson, and Jacob Carse. 2016 · 2016
Cited alongside, same era.
Towards end-to-end reinforcement learning of dialogue agents for information access
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomenech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al. 2017 · 2017
Later among the works it cites.
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2016 · 2016
Cited alongside, same era.
Policy networks with two-stage training for dialogue systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016 · 2016
Cited alongside, same era.
A user simulator for task-completion dialogues
Xiujun Li, Zachary C Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen. 2016 · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016 · 2016
Cited alongside, same era.
Continuously learning neural dialogue management
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
Tiancheng Zhao and Maxine Eskenazi. 2016 · 2016
Cited alongside, same era.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
End-to-end task-completion neural dialogue systems
Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017 · 2017
Cited alongside, same era.
Inigo Casanueva, Paweł Budzianowski, Pei-Hao Su, Stefan Ultes, Lina Rojas-Barahona, Bo-Hsiang Tseng, and Milica Gašić. 2018 · 2018
Later among the works it cites.
User modeling for task oriented dialogues
Izzeddin Gür, Dilek Hakkani-Tür, Gokhan Tür, and Pararth Shah. 2018 · 2018
Later among the works it cites.
Neural user simulation for corpus-based policy optimisation for spoken dialogue systems
Florian Kreyssig, Inigo Casanueva, Pawel Budzianowski, and Milica Gasic. 2018 · 2018
Later among the works it cites.
Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng. 2018 · 2018
Later among the works it cites.
Bing Liu, Gokhan Tur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018 · 2018
Later among the works it cites.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018 · 2018
Later among the works it cites.
Discriminative deep dyna-q: Robust planning for dialogue policy learning
Shang-Yu Su, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yun-Nung Chen. 2018 · 2018
Later among the works it cites.
Yuexin Wu, Xiujun Li, Jingjing Liu, Jianfeng Gao, and Yiming Yang. 2018 · 2018
Later among the works it cites.
Hierarchical text generation and planning for strategic dialogue
Denis Yarats and Mike Lewis. 2018 · 2018
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, Lihong Li, et al. 2019 · 2019
Later among the works it cites.