Fetching the paper…
Reading the bibliography…
Partially observable Markov decision processes (POMDPs) are a powerful abstraction for tasks that require decision making under uncertainty, and capture a wide range of real world tasks.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
The optimal control of partially observable markov processes over the infinite horizon: Discounted costs
Sondik, E. J. (1978) · 1978
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Singh, S. P., Jaakkola, T., and Jordan, M. I. (1994) · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G. (1995) · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Approximate planning in large pomdps via reusable trajectories
Kearns, M. J., Mansour, Y., and Ng, A. Y. (2000) · 2000
Earlier work this paper cites.
Reinforcement learning for spoken dialogue systems
Singh, S. P., Kearns, M. J., Litman, D. J., and Walker, M. A. (2000) · 2000
Earlier work this paper cites.
A (revised) survey of approximate methods for solving partially observable markov decision processes
Aberdeen, D. (2003) · 2003
Earlier work this paper cites.
Robot planning in partially observable continuous domains
Porta, J. M., Spaan, M. T., and Vlassis, N. (2005) · 2005
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Bayes-adaptive POMDPs
Ross, S., Chaib-draa, B., and Pineau, J. (2008) · 2008
Earlier work this paper cites.
Monte-Carlo planning in large POMDPs
Silver, D. and Veness, J. (2010) · 2010
Earlier work this paper cites.
Temporal-difference learning to assist human decision making during the control of an artificial limb
Edwards, A. L., Kearney, A., Dawson, M. R., Sutton, R. S., and Pilarski, P. M. (2013) · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Cited alongside, same era.
Amortized inference in probabilistic reasoning
Gershman, S. and Goodman, N. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
Prediction under uncertainty with error-encoding networks
Henaff, M., Zhao, J., and LeCun, Y. (2017) · 2017
Later among the works it cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J. (2017) · 2017
Later among the works it cites.
Learning in POMDPs with Monte Carlo tree search
Katt, S., Oliehoek, F. A., and Amato, C. (2017) · 2017
Later among the works it cites.
Learning to run challenge: Synthesizing physiologically accurate motion using deep reinforcement learning
Kidziński, Ł., Mohanty, S. P., Ong, C., Hicks, J., Francis, S., Levine, S., Salathé, M., and Delp, S. (2018) · 2017
Later among the works it cites.
Structured inference networks for nonlinear state space models
Krishnan, R. G., Shalit, U., and Sontag, D. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rezende, D. J., Mohamed, S., and Wierstra, D. (2014) · 2014
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. and Stone, P. (2015) · 2015
Cited alongside, same era.
Krishnan, R. G., Shalit, U., and Sontag, D. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. (2015) · 2015
Cited alongside, same era.
A deeper look at planning as learning from replay
Van Seijen, H. and Sutton, R. (2015) · 2015
Cited alongside, same era.
Sequential neural models with stochastic layers
Fraccaro, M., Sønderby, S. K., Paquet, U., and Winther, O. (2016) · 2016
Cited alongside, same era.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Learning multimodal transition dynamics for model-based reinforcement learning
Moerland, T. M., Broekens, J., and Jonker, C. M. (2017) · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al. (2017) · 2017
Later among the works it cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al. (2017) · 2017
Later among the works it cites.
Hybrid reward architecture for reinforcement learning
Van Seijen, H., Fatemi, M., Romoff, J., Laroche, R., Barnes, T., and Tsang, J. (2017) · 2017
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Buesing, L., Weber, T., Racaniere, S., Eslami, S., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., et al. (2018) · 2018
Closest in time.
Generative temporal models with spatial memory for partially observed environments
Fraccaro, M., Rezende, D. J., Zwols, Y., Pritzel, A., Eslami, S., and Viola, F. (2018) · 2018
Closest in time.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., and Levine, S. (2018) · 2018
Closest in time.