Fetching the paper…
Reading the bibliography…
Planning - the ability to analyze the structure of a problem in the large and decompose it into interrelated subproblems - is a hallmark of human intelligence.
A note on two problems in connexion with graphs
Dijkstra, E. W. et al · 1959
Earlier work this paper cites.
Experiments with the graph traverser program
Doran, J. E. and Michie, D · 1966
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B · 1968
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Navigation and acquisition of spatial knowledge in a virtual maze
Gillner, S. and Mallot, H. A · 1998
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning
LaValle, S. M · 1998
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Hallucinative topological memory for zero-shot visual planning
Liu, K., Kurutach, T., Tung, C. K.-C., Abbeel, P., and Tamar, A · 2002
Earlier work this paper cites.
Human spatial representation: Insights from animals
Wang, R. F. and Spelke, E. S · 2002
Earlier work this paper cites.
Plan2vec: Unsupervised representation learning by latent plans
Yang, G., Zhang, A., Morcos, A. S., Pineau, J., Abbeel, P., and Calandra, R · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Simultaneous localization and mapping: part i
Durrant-Whyte, H. and Bailey, T · 2006
Earlier work this paper cites.
k-means++: The advantages of careful seeding
Arthur, D. and Vassilvitskii, S · 2007
Earlier work this paper cites.
Robot navigation by waypoints
Wang, Y., Mulvaney, D., Sillitoe, I., and Swere, E · 2008
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
Treeqn and atreec: Differentiable tree-structured models for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2017
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
Later among the works it cites.
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C · 2018
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Eysenbach, B., Salakhutdinov, R. R., and Levine, S · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Later among the works it cites.
Mapping state space using landmarks for universal goal reaching
Huang, Z., Liu, F., and Su, H · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomènech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Hassabis, D., Silver, D., and Wierstra, D · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Efficient model-based deep reinforcement learning with variational state tabulation
Corneil, D., Gerstner, W., and Brea, J · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Cited alongside, same era.
Lee, L., Parisotto, E., Chaplot, D. S., Xing, E., and Salakhutdinov, R · 2018
Cited alongside, same era.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Planning with goal-conditioned policies
Nasiriany, S., Pong, V., Lin, S., and Levine, S · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
Zhao, R., Sun, X., and Tresp, V · 2019
Later among the works it cites.
Sparse graphical memory for robust planning
Emmons, S., Jain, A., Laskin, M., Kurutach, T., Abbeel, P., and Pathak, D · 2020
Closest in time.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Closest in time.
Sub-goal trees–a framework for goal-based reinforcement learning
Jurgenson, T., Avner, O., Groshev, E., and Tamar, A · 2020
Closest in time.
Long-horizon visual planning with goal-conditioned hierarchical predictors
Pertsch, K., Rybkin, O., Ebert, F., Finn, C., Jayaraman, D., and Levine, S · 2020
Closest in time.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Pitis, S., Chan, H., Zhao, S., Stadie, B., and Ba, J · 2020
Closest in time.