Fetching the paper…
Reading the bibliography…
Many studies have applied reinforcement learning to train a dialog policy and show great promise these years.
Learning to collaborate: Multi-scenario ranking via multi-agent reinforcement learning
Jun Feng, Heng Li, Minlie Huang, Shichen Liu, Wenwu Ou, Zhirong Wang, and Xiaoyan Zhu. 2018 · 1948
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson. 1983 · 1983
Earlier work this paper cites.
How to build user simulators to train rl-based dialog systems
Weiyan Shi, Kun Qian, Xuewei Wang, and Zhou Yu. 2019 · 2000
Earlier work this paper cites.
Dialogue act modeling for automatic tagging and recognition of conversational speech
Andreas Stolcke, Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, Daniel Jurafsky, Paul Taylor, Rachel Martin, Carol Van Ess-Dykema, and Marie Meteer. 2000 · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein. 2002 · 2002
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Parameter estimation for agenda-based user simulation
Simon Keizer, Milica Gašić, Filip Jurčíček, François Mairesse, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Earlier work this paper cites.
Single-agent vs. multi-agent techniques for concurrent reinforcement learning of negotiation dialogue policies
Kallirroi Georgila, Claire Nelson, and David Traum. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Earlier work this paper cites.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Layla El Asri, Jing He, and Kaheer Suleman. 2016 · 2016
Earlier work this paper cites.
Policy networks with two-stage training for dialogue systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016 · 2016
Earlier work this paper cites.
On-line active reward learning for policy optimisation in spoken dialogue systems
Pei-Hao Su, Milica Gašić, Nikola Mrkšić, Lina M Rojas Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Earlier work this paper cites.
The dialog state tracking challenge series: A review
Jason D Williams, Antoine Raux, and Matthew Henderson. 2016 · 2016
Cited alongside, same era.
Learning cooperative visual dialog agents with deep reinforcement learning
Abhishek Das, Satwik Kottur, José MF Moura, Stefan Lee, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Natural language does not emerge ‘naturally’ in multi-agent dialog
Satwik Kottur, José Moura, Stefan Lee, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018 · 2018
Later among the works it cites.
Neural user simulation for corpus-based policy optimisation of spoken dialogue systems
Florian Kreyssig, Iñigo Casanueva, Paweł Budzianowski, and Milica Gasic. 2018 · 2018
Later among the works it cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Later among the works it cites.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tür, Bing Liu, and Gokhan Tür. 2018 · 2018
Later among the works it cites.
Discriminative deep dyna-q: Robust planning for dialogue policy learning
Shang-Yu Su, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yun-Nung Chen. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George Van Den Driessche, Thore Graepel, and Demis Hassabis. 2017 · 2017
Cited alongside, same era.
Hybrid reward architecture for reinforcement learning
Harm Van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang. 2017 · 2017
Cited alongside, same era.
Multiwoz: A large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
User modeling for task oriented dialogues
Izzeddin Gür, Dilek Hakkani-Tür, Gokhan Tür, and Pararth Shah. 2018 · 2018
Cited alongside, same era.
Countering language drift via visual grounding
Jason Lee, Kyunghyun Cho, and Douwe Kiela. 2019a · 2019
Later among the works it cites.
Collaborative multi-agent dialogue model training via reinforcement learning
Alexandros Papangelis, Yi-Chia Wang, Piero Molino, and Gokhan Tur. 2019 · 2019
Later among the works it cites.
Guided dialog policy learning: Reward estimation for multi-domain task-oriented dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 2019
Later among the works it cites.
Budgeted policy learning for task-oriented dialogue systems
Zhirui Zhang, Xiujun Li, Jianfeng Gao, and Enhong Chen. 2019 · 2019
Later among the works it cites.
Rethinking action spaces for reinforcement learning in end-to-end dialog agents with latent variable models
Tiancheng Zhao, Kaige Xie, and Maxine Eskenazi. 2019 · 2019
Later among the works it cites.
Results of the multi-domain task-completion dialog challenge
Jinchao Li, Baolin Peng, Sungjin Lee, Jianfeng Gao, Ryuichi Takanobu, Qi Zhu, Minlie Huang, Hannes Schulz, Adam Atkinson, and Mahmoud Adada. 2020 · 2020
Closest in time.