Fetching the paper…
Reading the bibliography…
We introduce the $\gamma$-model, a predictive model of environment dynamics with an infinite probabilistic horizon.
Efficient Memory-based Learning for Robot Control
Moore, A. W · 1990
Earlier work this paper cites.
Neural networks for self-learning control systems
Nguyen, D. H. and Widrow, B · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
Jordan, M. I. and Rumelhart, D. E · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
Sutton, R · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1996
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Structure in the space of value functions
Foster, D. and Dayan, P · 2002
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
Silver, D., Sutton, R. S., and Müller, M · 2008
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Learning of closed-loop motion control
Farshidian, F., Neunert, M., and Buchli, J · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Approximate model-assisted neural fitted Q-iteration
Lampe, T. and Riedmiller, M · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Rezende, D. and Mohamed, S · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Deep successor reinforcement learning, 2016
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Cited alongside, same era.
Least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y. K., and Wang, Z · 2016
Cited alongside, same era.
f-gan: Training generative neural samplers using variational divergence minimization
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Later among the works it cites.
The successor representation: Its computational logic and neural substrates
Gershman, S. J · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nowozin, S., Cseke, B., and Tomioka, R · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J · 2017
Cited alongside, same era.
The successor representation in human reinforcement learning
Momennejad, I., Russek, E. M., Cheong, J. H., Botvinick, M. M., Daw, N. D., and Gershman, S. J · 2017
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Cited alongside, same era.
Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
Kahn, G., Villaflor, A., Ding, B., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Universal successor representations for transfer reinforcement learning
Ma, C., Wen, J., and Bengio, Y · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., S. Fearing, R., and Levine, S · 2018
Later among the works it cites.
Temporal difference models: Model-free deep RL for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
Later among the works it cites.
Combating the compounding-error problem with a multi-step model
Asadi, K., Misra, D., Kim, S., and Littman, M. L · 2019
Later among the works it cites.
Neural spline flows
Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2019
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Warde-Farley, D., de Wiele, T. V., and Mnih, V · 2020
Closest in time.