Fetching the paper…
Reading the bibliography…
In this paper, we present a deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
Ronald J Williams, · 1992
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Creating natural dialogs in the carnegie mellon communicator system.,”
Alexander I Rudnicky, Eric H Thayer, Paul C Constantinides, Chris Tchou, R Shern, Kevin A Lenzo, Wei Xu, and Alice Oh, · 1999
Earlier work this paper cites.
“Let’s go public! taking a spoken dialog system to the real world,”
Antoine Raux, Brian Langner, Dan Bohus, Alan W Black, and Maxine Eskenazi, · 2005
Earlier work this paper cites.
“Learning user simulations for information state update dialogue systems.,”
Kallirroi Georgila, James Henderson, and Oliver Lemon, · 2005
Earlier work this paper cites.
“Using pomdps for dialog management,”
Steve Young, · 2006
Earlier work this paper cites.
“A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies,”
Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young, · 2006
Earlier work this paper cites.
“Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets,”
James Henderson, Oliver Lemon, and Kallirroi Georgila, · 2008
Earlier work this paper cites.
“Reinforcement learning for parameter estimation in statistical spoken dialogue systems,”
Filip Jurčíček, Blaise Thomson, and Steve Young, · 2012
Earlier work this paper cites.
“Pomdp-based statistical spoken dialog systems: A review,”
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams, · 2013
Earlier work this paper cites.
“On-line policy optimisation of bayesian spoken dialogue systems via human interaction,”
M Gašić, Catherine Breslin, Matthew Henderson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young, · 2013
Earlier work this paper cites.
“Structured discriminative model for dialog state tracking,”
Sungjin Lee, · 2013
Earlier work this paper cites.
“Word-based dialog state tracking with recurrent neural networks,”
Matthew Henderson, Blaise Thomson, and Steve Young, · 2014
Earlier work this paper cites.
“Gaussian processes for pomdp-based dialogue manager optimization,”
Milica Gasic and Steve Young, · 2014
Earlier work this paper cites.
“Co-adaptation in spoken dialogue systems,”
Senthilkumar Chandramohan, Matthieu Geist, Fabrice Lefevre, and Olivier Pietquin, · 2014
Earlier work this paper cites.
“Single-agent vs. multi-agent techniques for concurrent reinforcement learning of negotiation dialogue policies.,”
Kallirroi Georgila, Claire Nelson, and David R Traum, · 2014
Cited alongside, same era.
“The second dialog state tracking challenge,”
Matthew Henderson, Blaise Thomson, and Jason Williams, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting.,”
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Neural responding machine for short-text conversation,”
Lifeng Shang, Zhengdong Lu, and Hang Li, · 2015
Cited alongside, same era.
“Using recurrent neural networks for slot filling in spoken language understanding,”
“A sequence-to-sequence model for user simulation in spoken dialogue systems,”
Layla El Asri, Jing He, and Kaheer Suleman, · 2016
Later among the works it cites.
“A user simulator for task-completion dialogues,”
Xiujun Li, Zachary C Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen, · 2016
Later among the works it cites.
“Deep reinforcement learning for dialogue generation,”
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky, · 2016
Later among the works it cites.
“End-to-end lstm-based dialog control optimized with supervised and reinforcement learning,”
Jason D Williams and Geoffrey Zweig, · 2016
Later among the works it cites.
“The dialog state tracking challenge series: A review,”
Jason Williams, Antoine Raux, and Matthew Henderson, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grégoire Mesnil, Yann Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng, Dilek Hakkani-Tur, Xiaodong He, Larry Heck, Gokhan Tur, Dong Yu, et al., · 2015
Cited alongside, same era.
“A neural conversational model,”
Oriol Vinyals and Quoc Le, · 2015
Cited alongside, same era.
“Human-machine dialogue as a stochastic game,”
Merwan Barlier, Julien Perolat, Romain Laroche, and Olivier Pietquin, · 2015
Cited alongside, same era.
“Machine learning for dialog state tracking: A review,”
Matthew Henderson, · 2015
Cited alongside, same era.
“Building end-to-end dialogue systems using generative hierarchical neural network models,”
Iulian V Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau, · 2016
Cited alongside, same era.
“A persona-based neural conversation model,”
Jiwei Li, Michel Galley, Chris Brockett, Georgios P Spithourakis, Jianfeng Gao, and Bill Dolan, · 2016
Cited alongside, same era.
“End-to-end memory networks with knowledge carryover for multi-turn spoken language understanding.,”
Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao, and Li Deng, · 2016
Cited alongside, same era.
“Conditional generation and snapshot learning in neural dialogue systems,”
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina M Rojas-Barahona, Pei-Hao Su, Stefan Ultes, David Vandyke, and Steve Young, · 2016
Later among the works it cites.
“An end-to-end trainable neural network model with belief tracking for task-oriented dialog,”
Bing Liu and Ian Lane, · 2017
Closest in time.
“Neural belief tracker: Data-driven dialogue state tracking,”
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young, · 2017
Closest in time.
“Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management,”
Pei-Hao Su, Pawel Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young, · 2017
Closest in time.
“A network-based end-to-end trainable task-oriented dialogue system,”
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young, · 2017
Closest in time.
“An end-to-end trainable neural network model with belief tracking for task-oriented dialog,”
Bing Liu and Ian Lane, · 2017
Closest in time.
“End-to-end task-completion neural dialogue systems,”
Xuijun Li, Yun-Nung Chen, Lihong Li, and Jianfeng Gao, · 2017
Closest in time.
“Learning end-to-end goal-oriented dialog,”
Antoine Bordes and Jason Weston, · 2017
Closest in time.
“Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning,”
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig, · 2017
Closest in time.