Fetching the paper…
Reading the bibliography…
Sequential decision making, commonly formalized as optimization of a Markov Decision Process, is a key challenge in artificial intelligence.
Benchmarking Model-Based Reinforcement Learning
Wang, T., Bao, X., Clavera, I., Hoang, J., Wen, Y., Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J. (2019) · 1907
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
The theory of dynamic programming
Bellman, R. (1954) · 1954
Earlier work this paper cites.
A Markovian decision process
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Heuristic problem solving: The next advance in operations research
Simon, H. A. and Newell, A. (1958) · 1958
Earlier work this paper cites.
A note on two problems in connexion with graphs
Dijkstra, E. W. (1959) · 1959
Earlier work this paper cites.
The shortest path through a maze
Moore, E. F. (1959) · 1959
Earlier work this paper cites.
Dynamic programming and markov processes
Howard, R. A. (1960) · 1960
Earlier work this paper cites.
Dynamic programming
Bellman, R. (1966) · 1966
Earlier work this paper cites.
Some studies in machine learning using the game of checkers. II - Recent progress
Samuel, A. L. (1967) · 1967
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B. (1968) · 1968
Earlier work this paper cites.
Heuristic search viewed as path finding in a graph
Pohl, I. (1970) · 1970
Earlier work this paper cites.
Problem-solving methods in Artificial Intelligence
Nilsson, N. J. (1971) · 1971
Earlier work this paper cites.
Depth-first search and linear graph algorithms
Tarjan, R. (1972) · 1972
Earlier work this paper cites.
Binary decision diagrams
Akers, S. B. (1978) · 1978
Earlier work this paper cites.
Planning and acting
McDermott, D. (1978) · 1978
Earlier work this paper cites.
Principles of artificial intelligence
Nilsson, N. J. (1982) · 1982
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W. (1983) · 1983
Earlier work this paper cites.
Chess 4.5—the Northwestern University chess program
Slate, D. J. and Atkin, L. R. (1983) · 1983
Earlier work this paper cites.
A multiple shooting algorithm for direct solution of optimal control problems
Bock, H. G. and Plitt, K.-J. (1984) · 1984
Earlier work this paper cites.
Heuristics: intelligent search strategies for computer problem solving
Pearl, J. (1984) · 1984
Earlier work this paper cites.
Depth-first iterative-deepening: An optimal admissible tree search
Korf, R. E. (1985) · 1985
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Real-time heuristic search
Korf, R. E. (1990) · 1990
Earlier work this paper cites.
Receding horizon control of nonlinear systems
Mayne, D. Q. and Michalska, H. (1990) · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
An analysis of stochastic shortest path problems
Bertsekas, D. P. and Tsitsiklis, J. N. (1991) · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J. (1991) · 1991
Earlier work this paper cites.
Symbolic boolean manipulation with ordered binary-decision diagrams
Bryant, R. E. (1992) · 1992
Earlier work this paper cites.
Planning as Satisfiability
Kautz, H. A., Selman, B., et al. (1992) · 1992
Earlier work this paper cites.
Efficient Memory-Bounded Search Methods
Russell, S. J. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Learning in embedded systems
Kaelbling, L. P. (1993) · 1993
Earlier work this paper cites.
Linear-space best-first search
Korf, R. E. (1993) · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W. and Atkeson, C. G. (1993) · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, A. G., Bradtke, S. J., and Singh, S. P. (1995) · 1995
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D. P. (1995) · 1995
Earlier work this paper cites.
Limited discrepancy search
Harvey, W. D. and Ginsberg, M. L. (1995) · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, S. P. and Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
On-line Policy Improvement using Monte-Carlo Search
Tesauro, G. and Galperin, G. R. (1997) · 1997
Earlier work this paper cites.
Bayesian Q-learning
Dearden, R., Friedman, N., and Russell, S. (1998) · 1998
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning
LaValle, S. M. (1998) · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J. (1999) · 1999
Earlier work this paper cites.
Model predictive control: past, present and future
Morari, M. and Lee, J. H. (1999) · 1999
Earlier work this paper cites.
Evolutionary algorithms for reinforcement learning
Moriarty, D. E., Schultz, A. C., and Grefenstette, J. J. (1999) · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Planning as heuristic search
Bonet, B. and Geffner, H. (2001) · 2001
Earlier work this paper cites.
LAO ⋆ \star : A heuristic search algorithm that finds solutions with loops
Hansen, E. A. and Zilberstein, S. (2001) · 2001
Earlier work this paper cites.
The FF planning system: Fast plan generation through heuristic search
Hoffmann, J. and Nebel, B. (2001) · 2001
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Cited alongside, same era.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Cited alongside, same era.
Deep blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h. (2002) · 2002
Cited alongside, same era.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Kearns, M., Mansour, Y., and Ng, A. Y. (2002) · 2002
Cited alongside, same era.
A reinforcement learning method based on adaptive simulated annealing
Atiya, A. F., Parlos, A. G., and Ingber, L. (2003) · 2003
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Puterman, M. L. (2014) · 2014
Later among the works it cites.
Balancing exploration and exploitation in classical planning
Schulte, T. and Keller, T. (2014) · 2014
Later among the works it cites.
Deterministic Policy Gradient Algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Later among the works it cites.
A comparison of knowledge-based GBFS enhancements and knowledge-free exploration
Valenzano, R. A., Sturtevant, N. R., Schaeffer, J., and Xie, F. (2014) · 2014
Later among the works it cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y. (2015) · 2015
Later among the works it cites.
Anytime optimal MDP planning with trial-based heuristic tree search
Keller, T. (2015) · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S. (2003) · 2003
Cited alongside, same era.
KBFS: K-best-first search
Felner, A., Kraus, S., and Korf, R. E. (2003) · 2003
Cited alongside, same era.
The cross entropy method for fast policy search
Mannor, S., Rubinstein, R. Y., and Gat, Y. (2003) · 2003
Cited alongside, same era.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A. G., and Singh, S. P. (2005) · 2005
Cited alongside, same era.
Bounded real-time dynamic programming: RTDP with monotone upper bounds and performance guarantees
McMahan, H. B., Likhachev, M., and Gordon, G. J. (2005) · 2005
Cited alongside, same era.
Think Too Fast Nor Too Slow: The Computational Trade-off Between Planning And Reinforcement Learning
Moerland, T. M., Deichler, A., Baldi, S., Broekens, J., and Jonker, C. M. (2020b) · 2005
Cited alongside, same era.
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Later among the works it cites.
Blundell, C., Uria, B., Pritzel, A., Li, Y., Ruderman, A., Leibo, J. Z., Rae, J., Wierstra, D., and Hassabis, D. (2016) · 2016
Later among the works it cites.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Later among the works it cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al. (2016) · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J. (2016) · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. (2016) · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Later among the works it cites.
Artificial intelligence: a modern approach
Russell, S. J. and Norvig, P. (2016) · 2016
Later among the works it cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, J. and Sastry, S. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A. (2017) · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Later among the works it cites.
Boltzmann exploration done right
Cesa-Bianchi, N., Gentile, C., Lugosi, G., and Neu, G. (2017) · 2017
Later among the works it cites.
Reinforcement learning and episodic memory in humans and animals: an integrative framework
Gershman, S. J. and Daw, N. D. (2017) · 2017
Later among the works it cites.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C. (2017) · 2017
Later among the works it cites.
Best-first width search: Exploration and exploitation in classical planning
Lipovetzky, N. and Geffner, H. (2017) · 2017
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2017) · 2017
Later among the works it cites.
Efficient exploration with double uncertain value networks
Moerland, T. M., Broekens, J., and Jonker, C. M. (2017) · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Later among the works it cites.
Neural Episodic Control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C. (2017) · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
Scalable planning with tensorflow for hybrid nonlinear domains
Wu, G., Say, B., and Sanner, S. (2017) · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H. (2018) · 2018
Later among the works it cites.
Efficient model-based deep reinforcement learning with variational state tabulation
Corneil, D., Gerstner, W., and Brea, J. (2018) · 2018
Later among the works it cites.
Forward-backward reinforcement learning
Edwards, A. D., Downs, L., and Davidson, J. C. (2018) · 2018
Later among the works it cites.
Automatic Goal Generation for Reinforcement Learning Agents
Florensa, C., Held, D., Geng, X., and Abbeel, P. (2018) · 2018
Later among the works it cites.
An introduction to deep reinforcement learning
François-Lavet, V., Henderson, P., Islam, R., Bellemare, M. G., Pineau, J., et al. (2018) · 2018
Later among the works it cites.
The Control Handbook (three volume set)
Levine, W. S. (2018) · 2018
Later among the works it cites.
The Potential of the Return Distribution for Exploration in RL
Moerland, T. M., Broekens, J., and Jonker, C. M. (2018) · 2018
Later among the works it cites.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Péré, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y. (2018) · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. (2018) · 2018
Later among the works it cites.
Solving the Rubik’s cube with deep reinforcement learning and search
Agostinelli, F., McAleer, S., Shmakov, A., and Baldi, P. (2019) · 2019
Later among the works it cites.
Analogues of mental simulation and imagination in deep learning
Hamrick, J. B. (2019) · 2019
Later among the works it cites.
Bootstrapping upper confidence bound
Hao, B., Abbasi Yadkori, Y., Wen, Z., and Cheng, G. (2019) · 2019
Later among the works it cites.
Introduction to Multi-Armed Bandits
Slivkins, A. et al. (2019) · 2019
Later among the works it cites.
Combining q-learning and search with amortized value estimates
Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Pfaff, T., Weber, T., Buesing, L., and Battaglia, P. W. (2020) · 2020
Closest in time.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D. (2020) · 2020
Closest in time.
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. (2021) · 2021
Closest in time.
High-accuracy model-based reinforcement learning, a survey
Plaat, A., Kosters, W., and Preuss, M. (2021) · 2021
Closest in time.