Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) algorithms can use high-capacity deep networks to learn directly from image observations.
Optimal control of markov processes with incomplete state information
K. J. Astrom · 1965
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
M. L. Littman, A. R. Cassandra, and L. P. Kaelbling · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
S. Lange and M. Riedmiller · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Black box variational inference for state space models
E. Archer, I. M. Park, L. Buesing, J. Cunningham, and L. Paninski · 2015
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
M. Hausknecht and P. Stone · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
R. G. Krishnan, U. Shalit, and D. Sontag · 2015
Earlier work this paper cites.
From pixels to torques: Policy learning with deep dynamical models
N. Wahlström, T. B. Schön, and M. P. Deisenroth · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
J. Foerster, I. A. Assael, N. de Freitas, and S. Whiteson · 2016
Cited alongside, same era.
Sequential neural models with stochastic layers
M. Fraccaro, S. K. Sonderby, U. Paquet, and O. Winther · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell · 2016
Cited alongside, same era.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Cited alongside, same era.
A disentangled recognition and nonlinear dynamics model for unsupervised learning
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Later among the works it cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Fraccaro, S. Kamronn, U. Paquet, and O. Winther · 2017
Cited alongside, same era.
DARLA: Improving zero-shot transfer in reinforcement learning
I. Higgins, A. Pal, A. Rusu, L. Matthey, C. Burgess, A. Pritzel, M. Botvinick, C. Blundell, and A. Lerchner · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Fixing a broken elbo
A. Alemi, B. Poole, I. Fischer, J. Dillon, R. A. Saurous, and K. Murphy · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Learning and querying fast generative models for reinforcement learning
L. Buesing, T. Weber, S. Racanière, S. M. A. Eslami, D. J. Rezende, D. P. Reichert, F. Viola, F. Besse, K. Gregor, D. Hassabis, and D. Wierstra · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Later among the works it cites.
On improving deep reinforcement learning for POMDPs
P. Zhu, X. Li, P. Poupart, and G. Miao · 2018
Later among the works it cites.
The value function polytope in reinforcement learning
R. Dadashi, A. A. Taïga, N. L. Roux, D. Schuurmans, and M. G. Bellemare · 2019
Closest in time.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Closest in time.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Closest in time.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Closest in time.
Variational temporal abstraction
T. Kim, S. Ahn, and Y. Bengio · 2019
Closest in time.
Biva: A very deep hierarchy of latent variables for generative modeling
L. Maaloe, M. Fraccaro, V. Liévin, and O. Winther · 2019
Closest in time.
Generating diverse high-fidelity images with VQ-VAE-2
A. Razavi, A. v. d. Oord, and O. Vinyals · 2019
Closest in time.
SOLAR: Deep structured latent representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. J. Johnson, and S. Levine · 2019
Closest in time.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Closest in time.