Fetching the paper…
Reading the bibliography…
Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumption: The order in which a plan is constructed is the same in which it is executed.
Principles of Artificial Intelligence
Nilsson, N. J · 1980
Earlier work this paper cites.
Taking Advantage of Stable Sets of Variables in Constraint Satisfaction Problems
Freuder, E. C. and Quinn, M. J · 1985
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P., Bertsekas, D. P., Bertsekas, D. P., and Bertsekas, D. P · 1995
Earlier work this paper cites.
Faster shortest-path algorithms for planar graphs
Henzinger, M. R., Klein, P., Rao, S., and Subramanian, S · 1997
Earlier work this paper cites.
The divide-and-conquer subgoal-ordering algorithm for speeding up logic inference
Ledeniov, O. and Markovitch, S · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Pseudo-tree search with soft constraints
Larrosa, J., Meseguer, P., and Sánchez, M · 2002
Earlier work this paper cites.
Hybrid backtracking bounded by tree-decomposition of constraint networks
Jégou, P. and Terrioux, C · 2003
Earlier work this paper cites.
Mixtures of deterministic-probabilistic networks and their AND/OR search space
Dechter, R. and Mateescu, R · 2004
Earlier work this paper cites.
AND/OR tree search for constraint optimization
Marinescu, R. and Dechter, R · 2004
Earlier work this paper cites.
Probabilistic planning vs. replanning
Little, I. and Thiébaux, S · 2007
Earlier work this paper cites.
Multi-armed bandits with episode context
Rosin, C. D · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J · 2016
Cited alongside, same era.
Nowak-Vila, A., Folqué, D., and Bruna, J · 2016
Floyd-Warshall Reinforcement Learning Learning from Past Experiences to Reach New Goals
Dhiman, V., Banerjee, S., Siskind, J. M., and Corso, J. J · 2018
Later among the works it cites.
Learning Actionable Representations with Goal-Conditioned Policies
Ghosh, D., Gupta, A., and Levine, S · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
Savinov, N., Dosovitskiy, A., and Koltun, V · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2018
Later among the works it cites.
Learning over Subgoals for Efficient Navigation of Structured, Unknown Environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Hindsight Experience Replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Cited alongside, same era.
Intention-net: Integrating planning and deep learning for goal-directed autonomous navigation
Gao, W., Hsu, D., Lee, W. S., Shen, S., and Subramanian, K · 2017
Cited alongside, same era.
Learning composable models of parameterized skills
Kaelbling, L. P. and Lozano-Pérez, T · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Cited alongside, same era.
Stein, G. J., Bradley, C., and Roy, N · 2018
Later among the works it cites.
Composable planning with attributes
Zhang, A., Lerer, A., Sukhbaatar, S., Fergus, R., and Szlam, A · 2018
Later among the works it cites.
Subgoal-based temporal abstraction in Monte-Carlo tree search
Gabor, T., Peter, J., Phan, T., Meyer, C., and Linnhoff-Popien, C · 2019
Later among the works it cites.
Sub-Goal Trees–a Framework for Goal-Directed Trajectory Prediction and Optimization
Jurgenson, T., Groshev, E., and Tamar, A · 2019
Later among the works it cites.
Depth-First Proof-Number Search with Heuristic Edge Cost and Application to Chemical Synthesis Planning
Kishimoto, A., Buesser, B., Chen, B., and Botea, A · 2019
Later among the works it cites.
Hierarchical visuomotor control of humanoids
Merel, J., Ahuja, A., Pham, V., Tunyasuvunakool, S., Liu, S., Tirumala, D., Heess, N., and Wayne, G · 2019
Later among the works it cites.
Planning with Goal-Conditioned Policies
Nasiriany, S., Pong, V., Lin, S., and Levine, S · 2019
Later among the works it cites.
Combining Q-Learning and Search with Amortized Value Estimates
Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Pfaff, T., Weber, T., Buesing, L., and Battaglia, P. W · 2020
Closest in time.