Fetching the paper…
Reading the bibliography…
Learned world models summarize an agent's experience to facilitate learning complex behaviors.
A new approach to linear filtering and prediction problems
R. E. Kalman · 1960
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
An introduction to variational methods for graphical models
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul · 1999
Earlier work this paper cites.
The information bottleneck method
N. Tishby, F. C. Pereira, and W. Bialek · 2000
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
R. G. Krishnan, U. Shalit, and D. Sontag · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
Deep variational information bottleneck
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy · 2016
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, et al · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
M. Karl, M. Soelch, J. Bayer, and P. van der Smagt · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Robust locally-linear controllable embedding
E. Banijamali, R. Shu, M. Ghavamzadeh, H. Bui, and A. Ghodsi · 2017
Cited alongside, same era.
J. V. Dillon, I. Langmore, D. Tran, E. Brevdo, S. Vasudevan, D. Moore, B. Patton, A. Alemi, M. Hoffman, and R. A. Saurous · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Model-based planning with discrete and continuous actions
M. Henaff, W. F. Whitney, and Y. LeCun · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Ebert, C. Finn, A. X. Lee, and S. Levine · 2017
Cited alongside, same era.
Model-based planning in discrete action spaces
M. Henaff, W. F. Whitney, and Y. LeCun · 2017
Cited alongside, same era.
Value prediction network
J. Oh, S. Singh, and H. Lee · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Later among the works it cites.
Formal limitations on the measurement of mutual information
D. McAllester and K. Statos · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Later among the works it cites.
Learning real-world robot policies by dreaming
A. Piergiovanni, A. Wu, and M. S. Ryoo · 2018
Later among the works it cites.
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Later among the works it cites.
Solar: deep structured representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine · 2018
Later among the works it cites.
Imagined value gradients: Model-based policy optimization with transferable latent dynamics models
A. Byravan, J. T. Springenberg, A. Abdolmaleki, R. Hafner, M. Neunert, T. Lampe, N. Siegel, N. Heess, and M. Riedmiller · 2019
Closest in time.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Closest in time.
Shaping belief states with generative environment models for rl
K. Gregor, D. J. Rezende, F. Besse, Y. Wu, H. Merzic, and A. v. d. Oord · 2019
Closest in time.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
M. Henaff, A. Canziani, and Y. LeCun · 2019
Closest in time.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, et al · 2019
Closest in time.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2019
Closest in time.
Pipps: Flexible model-based policy search robust to the curse of chaos
P. Parmas, C. E. Rasmussen, J. Peters, and K. Doya · 2019
Closest in time.
On variational bounds of mutual information
B. Poole, S. Ozair, A. v. d. Oord, A. A. Alemi, and G. Tucker · 2019
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2019
Closest in time.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Closest in time.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Closest in time.