Fetching the paper…
Reading the bibliography…
Despite recent advances in natural language understanding and generation, and decades of research on the development of conversational bots, building automated agents that can carry on rich open-ended conversations with humans "in the wild" remains a formidable challenge.
Asynchronous methods for deep reinforcement learning. In International Conf. on machine learning . PMLR, 1928–1937
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. 1991 · 1991
Earlier work this paper cites.
Varieties of Helmholtz machine
P. Dayan and G. E. Hinton. 1996 · 1996
Earlier work this paper cites.
Learning dialogue strategies within the Markov decision process framework. In 1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proc. 72–79
E. Levin, R. Pieraccini, and W. Eckert. 1997 · 1997
Earlier work this paper cites.
Reinforcement Learning for Spoken Dialogue Systems. In Proc. of the 12th International Conf. on Neural Information Processing Systems (NIPS’99) . MIT Press, 956–962
S. Singh, M. Kearns, D. Litman, and M. Walker. 1999 · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour. 1999 · 1999
Earlier work this paper cites.
An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email
M. A. Walker. 2000 · 2000
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the NJFun system
S. Singh, D. Litman, M. Kearns, and M. Walker. 2002 · 2002
Earlier work this paper cites.
Mixture of expert agents for handling imbalanced data sets
S. Kotsiantis and P. Pintelas. 2003 · 2003
Earlier work this paper cites.
Sentence fusion for multidocument news summarization
R. Barzilay and K. R. McKeown. 2005 · 2005
Earlier work this paper cites.
Explorations in Sentence Fusion. In Proc. of the Tenth European Workshop on Natural Language Generation
E. Marsi and E. Krahmer. 2005 · 2005
Earlier work this paper cites.
Partially observable Markov decision processes for spoken dialog systems
J. D. Williams and S. Young. 2007 · 2007
Earlier work this paper cites.
Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets
J. Henderson, O. Lemon, and K. Georgila. 2008 · 2008
Earlier work this paper cites.
Reinforcement learning for dialog management using least-squares policy iteration and fast feature selection. In 10th Annual Conf. of the International Speech Communication Association
L. Li, J. D. Williams, and S. Balakrishnan. 2009 · 2009
Earlier work this paper cites.
Sparse approximate dynamic programming for dialog management. In Proc. of the SIGDIAL 2010 Conf. 107–115
S. Chandramohan, M. Geist, and O. Pietquin. 2010 · 2010
Earlier work this paper cites.
The hidden information state model: A practical framework for POMDP-based spoken dialogue management
S. Young, M. Gašić, S. Keizer, F. Mairesse, J. Schatzmann, B. Thomson, and K. Yu. 2010 · 2010
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects. In IEEE Workshop on Automatic Speech Recognition & Understanding . 312–317
M. Gašić, F. Jurčíček, B. Thomson, K. Yu, and S. Young. 2011 · 2011
Cited alongside, same era.
Policy search for motor primitives in robotics
J. Kober and J. Peters. 2011 · 2011
Cited alongside, same era.
An unbiased offline evaluation of contextual bandit algorithms with generalized linear models. In Proc. of the Workshop on On-line Trading of Exploration and Exploitation 2 . 19–36
L. Li, W. Chu, J. Langford, T. Moon, and X. Wang. 2012 · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller. 2013 · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms. In Conf. on machine learning
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller. 2014 · 2014
D. Cer, Y. Yang, S. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, Y. Sung, B. Strope, and R. Kurzweil. 2018 · 2018
Later among the works it cites.
B. Liu, G. Tur, D. Hakkani-Tur, P. Shah, and L. Heck. 2018 · 2018
Later among the works it cites.
DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion. In Proc. of the 2019 Conf. of the North American Chapter of the Association for Comp. Linguistics: Human Language Technologies . 3443–3455
M. Geva, E. Malmi, I. Szpektor, and J. Berant. 2019 · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-confidence off-policy evaluation. In Proc. of the AAAI Conf. on Artificial Intelligence , Vol. 29
P. Thomas, G. Theocharous, and M. Ghavamzadeh. 2015 · 2015
Cited alongside, same era.
Sample-efficient deep reinforcement learning for dialog control
K. Asadi and J. D. Williams. 2016 · 2016
Cited alongside, same era.
Policy networks with two-stage training for dialogue systems
M. Fatemi, L. E. Asri, H. Schulz, J. He, and K. Suleman. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
J. Li, W. Monroe, A. Ritter, M. Galley, J. Gao, and D. Jurafsky. 2016 · 2016
Cited alongside, same era.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, S. Jean, A. Ritter, and D. Jurafsky. 2017 · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
J. Schulman, X. Chen, and P. Abbeel. 2017 · 2017
Cited alongside, same era.
A Deep Reinforcement Learning Chatbot
I. V. Serban, C. Sankar, M. Germain, S. Zhang, Z. Lin, S. Subramanian, T. Kim, M. Pieper, S. Chandar, N. R. Ke, S. Mudumba, A. de Brébisson, J. Sotelo, D. Suhubdy, V. Michalski, A. Nguyen, J. Pineau, and Y. Bengio. 2017 · 2017
Cited alongside, same era.
Encode, Tag, Realize: High-Precision Text Editing. In Proc. of the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th International Joint Conf. on Natural Language Processing . Association for Comp. Linguistics, 5053–5064
E. Malmi, S. Krause, S. Rothe, D. Mirylenka, and A. Severyn. 2019 · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li. 2019 · 2019
Later among the works it cites.
CAQL: Continuous Action Q-Learning. In International Conf. on Learning Representations
M. Ryu, Y. Chow, R. Anderson, C. Tjandraatmadja, and C. Boutilier. 2019 · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum. 2019 · 2019
Later among the works it cites.
Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models. In Proc. of the 2019 Conf. of the North American Chapter of the Association for Comp. Linguistics: Human Language Technologies . 1208–1218
T. Zhao, K. Xie, and M. Eskenazi. 2019 · 2019
Later among the works it cites.
Human-centric dialog training via offline reinforcement learning
N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. S. Gu, and R. Picard. 2020 · 2020
Later among the works it cites.
Conservative Q-Learning for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 33. Curran Associates, Inc., 1179–1191
A. Kumar, A. Zhou, G. Tucker, and S. Levine. 2020 · 2020
Later among the works it cites.
Hierarchical reinforcement learning for open-domain dialog. In Proc. of the AAAI Conf. on Artificial Intelligence , Vol. 34. 8741–8748
A. Saleh, N. Jaques, A. Ghandeharioun, J. Shen, and R. Picard. 2020 · 2020
Later among the works it cites.
Generating empathetic responses by looking ahead the user’s sentiment. In IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP) . 7989–7993
J. Shin, P. Xu, A. Madotto, and P. Fung. 2020 · 2020
Later among the works it cites.
Dynamic composition for conversational domain exploration. In Proc. of The Web Conf. 2020 . 872–883
I. Szpektor, D. Cohen, G. Elidan, M. Fink, A. Hassidim, O. Keller, S. Kulkarni, E. Ofek, S. Pudinsky, A. Revach, et al · 2020
Later among the works it cites.
The design and implementation of xiaoice, an empathetic social chatbot
L. Zhou, J. Gao, D. Li, and H.-Y. Shum. 2020 · 2020
Later among the works it cites.