Fetching the paper…
Reading the bibliography…
In this work, we propose an adversarial learning method for reward estimation in reinforcement learning (RL) based task-oriented dialog models.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Paradise: A framework for evaluating spoken dialogue agents
Marilyn A Walker, Diane J Litman, Candace A Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al. 2000 · 2000
Earlier work this paper cites.
Automating spoken dialogue management design using machine learning: An industry perspective
Tim Paek and Roberto Pieraccini. 2008 · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell. 2010 · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
Predicting user satisfaction in spoken dialog system evaluation with collaborative filtering
Zhaojun Yang, Gina-Anne Levow, and Helen Meng. 2012 · 2012
Earlier work this paper cites.
On-line policy optimisation of bayesian spoken dialogue systems via human interaction
Milica Gašić, Catherine Breslin, Matthew Henderson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young. 2013 · 2013
Earlier work this paper cites.
Dialog state tracking challenge 2 & 3
Matthew Henderson, Blaise Thomson, and Jason Williams. 2013 · 2013
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Earlier work this paper cites.
Task completion transfer learning for reward inference
Layla El Asri, Romain Laroche, and Olivier Pietquin. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
The second dialog state tracking challenge
Matthew Henderson, Blaise Thomson, and Jason Williams. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Deep generative image models using a laplacian pyramid of adversarial networks
Emily L Denton, Soumith Chintala, Rob Fergus, et al. 2015 · 2015
Cited alongside, same era.
Learning from real users: Rating dialogue success with neural networks for reinforcement learning in spoken dialogue systems
Pei-Hao Su, David Vandyke, Milica Gasic, Dongho Kim, Nikola Mrksic, Tsung-Hsien Wen, and Steve Young. 2015 · 2015
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios P Spithourakis, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Learning end-to-end goal-oriented dialog
Antoine Bordes and Jason Weston. 2017 · 2017
Later among the works it cites.
Towards end-to-end reinforcement learning of dialogue agents for information access
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017 · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017b · 2017
Later among the works it cites.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Later among the works it cites.
Data-efficient deep reinforcement learning for dexterous manipulation
Ivaylo Popov, Nicolas Heess, Timothy Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerik, Thomas Lampe, Yuval Tassa, Tom Erez, and Martin Riedmiller. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural network models
Iulian V Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
On-line active reward learning for policy optimisation in spoken dialogue systems
Pei-Hao Su, Milica Gašić, Nikola Mrkšić, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
End-to-end lstm-based dialog control optimized with supervised and reinforcement learning
Jason D Williams and Geoffrey Zweig. 2016 · 2016
Cited alongside, same era.
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016 · 2016
Cited alongside, same era.
Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning
Tiancheng Zhao and Maxine Eskenazi. 2016 · 2016
Cited alongside, same era.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Alan Ritter, and Dan Jurafsky. 2017a
Cited in the paper.
Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management
Pei-Hao Su, Paweł Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young. 2017 · 2017
Later among the works it cites.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017 · 2017
Later among the works it cites.
Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Improving neural machine translation with conditional sequence generative adversarial nets
Zhen Yang, Wei Chen, Feng Wang, and Bo Xu. 2017 · 2017
Later among the works it cites.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gokhan Tur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018 · 2018
Closest in time.
Adversarial advantage actor-critic model for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Yun-Nung Chen, and Kam-Fai Wong. 2018 · 2018
Closest in time.