Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has recently shown many impressive successes.
Aggregation in dynamic programming
Bean, J. C., Birge, J. R., and Smith, R. L · 1987
Earlier work this paper cites.
Adaptive aggregation methods for infinite horizon dynamic programming
Bertsekas, D. P. and Castanon, D. A · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W. and Atkeson, C. G · 1993
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
Peng, J. and Williams, R. J · 1993
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, D. P. and Tsitsiklis, J. N · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Tesauro, G · 1995
Earlier work this paper cites.
Bagging predictors
Breiman, L · 1996
Earlier work this paper cites.
Model reduction techniques for computing approximately optimal solutions for Markov decision processes
Dean, T., Givan, R., and Leach, S · 1997
Earlier work this paper cites.
User modeling for spoken dialogue system evaluation
Eckert, W., Levin, E., and Pieraccini, R · 1997
Earlier work this paper cites.
Approximation in model-based learning
Kuvayev, L. and Sutton, R · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Efficient model-based exploration
Wiering, M. and Schmidhuber, J · 1998
Earlier work this paper cites.
Decision-theoretic planning: Structural assumptions and computational leverage
Boutilier, C., Dean, T., and Hanks, S · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Reinforcement learning for spoken dialogue systems
Singh, S. P., Kearns, M. J., Litman, D. J., and Walker, M. A · 1999
Earlier work this paper cites.
Reinforcement learning soccer teams with incomplete world models
Wiering, M., Sałustowicz, R., and Schmidhuber, J · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
A stochastic model of human-machine interaction for learning dialog strategies
Levin, E., Pieraccini, R., and Eckert, W · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Dialogue act modeling for automatic tagging and recognition of conversational speech
Stolcke, A., Ries, K., Coccaro, N., Shriberg, E., Bates, R., Jurafsky, D., Taylor, P., Martin, R., Van Ess-Dykema, C., and Meteer, M · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
Developing a flexible spoken dialog system using simulation
Chung, G · 2004
Earlier work this paper cites.
Metrics for finite Markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2004
Earlier work this paper cites.
Human-computer dialogue simulation using hidden Markov models
Cuayáhuitl, H., Renals, S., Lemon, O., and Shimodaira, H · 2005
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
Jong, N. K. and Stone, P · 2005
Earlier work this paper cites.
Learning the structure of factored markov decision processes in reinforcement learning problems
Degris, T., Sigaud, O., and Wuillemin, P.-H · 2006
Earlier work this paper cites.
User simulation for spoken dialogue systems: Learning and evaluation
Georgila, K., Henderson, J., and Lemon, O · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L · 2006
Cited alongside, same era.
Performance loss bounds for approximate value iteration with state aggregation
Van Roy, B · 2006
Cited alongside, same era.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Schatzmann, J., Thomson, B., Weilhammer, K., Ye, H., and Young, S · 2007
Cited alongside, same era.
Chatbots: are they really useful?
Shawar, B. A. and Atwell, E · 2007
Cited alongside, same era.
Efficient structure learning in factored-state mdps
Strehl, A. L., Diuk, C., and Littman, M. L · 2007
Cited alongside, same era.
Partially observable markov decision processes for spoken dialog systems
Williams, J. D. and Young, S · 2007
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Later among the works it cites.
Agnostic system identification for monte carlo planning
Talvitie, E · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Later among the works it cites.
Near optimal behavior via approximate state abstraction
Abel, D., Hershkowitz, D. E., and Littman, M. L · 2016
Later among the works it cites.
A sequence-to-sequence model for user simulation in spoken dialogue systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Model-based bayesian reinforcement learning in large structured domains
Ross, S. and Pineau, J · 2008
Cited alongside, same era.
Multi-party, multi-issue, multi-strategy negotiation for multi-modal virtual agents
Traum, D., Marsella, S. C., Gratch, J., Lee, J., and Hartholt, A · 2008
Cited alongside, same era.
Representing the reinforcement learning state in a negotiation dialogue
Heeman, P. A · 2009
Cited alongside, same era.
Are we there yet? research in commercial spoken dialog systems
Pieraccini, R., Suendermann, D., Dayanidhi, K., and Liscombe, J · 2009
Cited alongside, same era.
Introduction to nonparametric estimation. revised and extended from the 2004 french original. translated by vladimir zaiats, 2009
Tsybakov, A. B · 2009
Cited alongside, same era.
The anatomy of alice
Wallace, R. S · 2009
Cited alongside, same era.
Asri, L. E., He, J., and Suleman, K · 2016
Later among the works it cites.
Incremental stochastic factorization for online reinforcement learning
Barreto, A. d. M. S., Beirigo, R. L., Pineau, J., and Precup, D · 2016
Later among the works it cites.
Policy networks with two-stage training for dialogue systems
Fatemi, M., Asri, L. E., Schulz, H., He, J., and Suleman, K · 2016
Later among the works it cites.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Later among the works it cites.
Learning to query, reason, and answer questions on ambiguous texts
Guo, X., Klinger, T., Rosenbaum, C., Bigus, J. P., et al · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Liu, C.-W., Lowe, R., Serban, I. V., Noseworthy, M., Charlin, L., and Pineau, J · 2016
Later among the works it cites.
Automatic creation of scenarios for evaluating spoken dialogue systems via user-simulation
López-Cózar, R · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., et al · 2016
Later among the works it cites.
Continuously learning neural dialogue management
Su, P.-H., Gasic, M., Mrksic, N., Rojas-Barahona, L., et al · 2016
Later among the works it cites.
Improved learning of dynamics models for control
Venkatraman, A., Capobianco, R., Pinto, L., Hebert, M., Nardi, D., and Bagnell, J. A · 2016
Later among the works it cites.
Strategy and policy learning for non-task-oriented conversational systems
Yu, Z., Xu, Z., Black, A. W., and Rudnicky, A. I · 2016
Later among the works it cites.
Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning
Zhao, T. and Eskenazi, M · 2016
Later among the works it cites.
Goal-Driven Dynamics Learning via Bayesian Optimization
Bansal, S., Calandra, R., Xiao, T., Levine, S., and Tomlin, C. J · 2017
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
Das, A., Kottur, S., Moura, J. M., Lee, S., and Batra, D · 2017
Later among the works it cites.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S · 2017
Later among the works it cites.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., et al · 2017
Later among the works it cites.
Incremental human-machine dialogue simulation
Khouzaimi, H., Laroche, R., and Lefevre, F · 2017
Later among the works it cites.
Deal or No Deal? End-to-End Learning for Negotiation Dialogues
Lewis, M., Yarats, D., Dauphin, Y. N., Parikh, D., and Batra, D · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Liu, B. and Lane, I · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., et al · 2017
Later among the works it cites.
Conversational AI: The Science Behind the Alexa Prize
Ram, A., Prasad, R., Khatri, C., Venkatesh, A., et al · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., et al · 2017
Later among the works it cites.
Integrating planning for task-completion dialogue policy learning
Peng, B., Li, X., Gao, J., Liu, J., and Wong, K.-F · 2018
Closest in time.