Fetching the paper…
Reading the bibliography…
Learned dynamics models combined with both planning and policy learning algorithms have shown promise in enabling artificial agents to learn to perform many diverse tasks with limited supervision.
Unsupervised visuomotor control through distributional planning networks
T. Yu, G. Shevchuk, D. Sadigh, and C. Finn · 1902
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 1906
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. R. Julian, K. Hausman, C. Finn, and S. Levine · 1910
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 1912
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation and Machine Learning
R. Rubinstein and D. Kroese · 2004
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. T. Springenberg, J. Boedecker, and M. A. Riedmiller · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Earlier work this paper cites.
Learning to act by predicting the future
A. Dosovitskiy and V. Koltun · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. J. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Improving pilco with bayesian neural network dynamics models
R. McAllister and C. E. Rasmussen · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Robust locally-linear controllable embedding
E. Banijamali, R. Shu, M. Ghavamzadeh, H. H. Bui, and A. Ghodsi · 2017
Earlier work this paper cites.
End-to-end driving via conditional imitation learning
F. Codevilla, M. Müller, A. Dosovitskiy, A. López, and V. Koltun · 2017
Earlier work this paper cites.
Self-supervised visual planning with temporal skip connections
F. Ebert, C. Finn, A. X. Lee, and S. Levine · 2017
Earlier work this paper cites.
Value-Aware Loss Function for Model-based Reinforcement Learning
A.-M. Farahmand, A. Barreto, and D. Nikovski · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
S. Racanière, T. Weber, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. M. O. Heess, Y. Li, R. Pascanu, P. W. Battaglia, D. Hassabis, D. Silver, and D. Wierstra · 2017
Cited alongside, same era.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Stochastic video generation with a learned prior
E. L. Denton and R. Fergus · 2018
Cited alongside, same era.
Learning to predict without looking ahead: World models without forward prediction
D. Freeman, D. Ha, and L. Metz · 2019
Later among the works it cites.
Shaping belief states with generative environment models for rl
K. Gregor, D. J. Rezende, F. Besse, Y. Wu, H. Merzic, and A. van den Oord · 2019
Later among the works it cites.
Robot motion planning in learned latent spaces
B. Ichter and M. Pavone · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Ebert, C. Finn, S. Dasari, A. Xie, A. X. Lee, and S. Levine · 2018
Cited alongside, same era.
D. Ha and J. Schmidhuber · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Cited alongside, same era.
Learning plannable representations with causal infogan
T. Kurutach, A. Tamar, G. Yang, S. J. Russell, and P. Abbeel · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
A. Nagabandi, K. Konoglie, S. Levine, and V. Kumar · 2019
Later among the works it cites.
Contextual imagined goals for self-supervised robotic learning
A. Nair, S. Bahl, A. Khazatsky, V. Pong, G. Berseth, and S. Levine · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. P. Lillicrap, and D. Silver · 2019
Later among the works it cites.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks
B. Thananjeyan, A. Balakrishna, U. Rosolia, F. Li, R. McAllister, J. E. Gonzalez, S. Levine, F. Borrelli, and K. Goldberg · 2019
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
R. Veerapaneni, J. D. Co-Reyes, M. Chang, M. Janner, C. Finn, J. Wu, J. B. Tenenbaum, and S. Levine · 2019
Later among the works it cites.
High fidelity video prediction with large stochastic recurrent neural networks
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. Le, and H. Lee · 2019
Later among the works it cites.
Learning robotic manipulation through visual planning and acting
A. Wang, T. Kurutach, K. Liu, P. Abbeel, and A. Tamar · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Later among the works it cites.
Improvisation through physical understanding: Using novel objects as tools with visual foresight
A. Xie, F. Ebert, S. Levine, and C. Finn · 2019
Later among the works it cites.
Gradient-aware model-based policy search
P. D’Oro, A. M. Metelli, A. Tirinzoni, M. Papini, and M. Restelli · 2020
Closest in time.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
K. Hartikainen, X. Geng, T. Haarnoja, and S. Levine · 2020
Closest in time.
Learning latent state spaces for planning through reward prediction, 2020
A. Havens, Y. Ouyang, P. Nagarajan, and Y. Fujita · 2020
Closest in time.
Objective mismatch in model-based reinforcement learning
N. G. Lambert, B. Amos, O. Yadan, and R. Calandra · 2020
Closest in time.
Hallucinative topological memory for zero-shot visual planning, 2020
K. Liu, T. Kurutach, P. Abbeel, and A. Tamar · 2020
Closest in time.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
S. Nair and C. Finn · 2020
Closest in time.
Goal-conditioned video prediction, 2020
O. Rybkin, K. Pertsch, F. Ebert, D. Jayaraman, C. Finn, and S. Levine · 2020
Closest in time.
Model based reinforcement learning for atari
Łukasz Kaiser, M. Babaeizadeh, P. Miłos, B. Osiński, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker, and H. Michalewski · 2020
Closest in time.