Fetching the paper…
Reading the bibliography…
The history of learning for control has been an exciting back and forth between two broad classes of algorithms: planning and reinforcement learning.
Self-supervised learning of image embedding for continuous control
Florensa, C., Degrave, J., Heess, N., Springenberg, J. T., and Riedmiller, M. (2019) · 1901
Earlier work this paper cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P. (2019) · 1903
Earlier work this paper cites.
Rules for ordering uncertain prospects
Hadar, J. and Russell, W. R. (1969) · 1969
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Bickel, P. J., Freedman, D. A., et al. (1981) · 1981
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Probabilistic roadmaps for path planning in high-dimensional configuration spaces
Kavraki, L., Svestka, P., and Overmars, M. H. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J. (1998) · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, D. (2000) · 2000
Earlier work this paper cites.
Accelerating reinforcement learning by composing solutions of automatically identified subtasks
Drummond, C. (2002) · 2002
Earlier work this paper cites.
Principles of robot motion: theory, algorithms, and implementation
Choset, H. M., Hutchinson, S., Lynch, K. M., Kantor, G., Burgard, W., Kavraki, L. E., and Thrun, S. (2005) · 2005
Earlier work this paper cites.
Behavior planning for character animation
Lau, M. and Kuffner, J. J. (2005) · 2005
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Şimşek, Ö., Wolfe, A. P., and Barto, A. G. (2005) · 2005
Earlier work this paper cites.
Planning algorithms
LaValle, S. M. (2006) · 2006
Earlier work this paper cites.
Space-time planning with parameterized locomotion controllers
Levine, S., Lee, Y., Koltun, V., and Popović, Z. (2011) · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Learning navigation behaviors end-to-end with autorl
Chiang, H.-T. L., Faust, A., Fiser, M., and Francis, A. (2019) · 2014
Earlier work this paper cites.
Deepmpc: Learning deep latent features for model predictive control
Lenz, I., Knepper, R. A., and Saxena, A. (2015) · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. (2015) · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. (2015) · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A. V., Abbeel, P., Malik, J., and Levine, S. (2016) · 2016
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Semantic scene completion from a single depth image
Song, S., Yu, F., Zeng, A., Chang, A. X., Savva, M., and Funkhouser, T. (2017) · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., and Farhadi, A. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J. (2016) · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al. (2016) · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D. (2017) · 2017
Cited alongside, same era.
Differentiable mpc for end-to-end planning and control
Amos, B., Jimenez, I., Sacks, J., Boots, B., and Kolter, J. Z. (2018) · 2018
Later among the works it cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., and van den Hengel, A. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S. (2018) · 2018
Later among the works it cites.
Prm-rl: Long-range robotic navigation tasks by combining reinforcement learning and sampling-based planning
Faust, A., Ramirez, O., Fiser, M., Oslund, K., Francis, A., Davidson, J., and Tapia, L. (2018) · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Lee, L., Parisotto, E., Chaplot, D. S., Xing, E., and Salakhutdinov, R. (2018) · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S. (2018) · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S. (2018) · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Later among the works it cites.
Shah, P., Fiser, M., Faust, A., Kew, J. C., and Hakkani-Tur, D. (2018) · 2018
Later among the works it cites.
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
The laplacian in rl: Learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O. (2018) · 2018
Later among the works it cites.
Hierarchical reinforcement learning with hindsight
Levy, A., Platt, R., and Saenko, K. (2019) · 2019
Closest in time.