Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) is a sample-efficient way of learning complex behaviors by leveraging a learned single-step dynamics model to plan actions in imagination.
Temporal credit assignment in reinforcement learning
R. S. Sutton · 1984
Earlier work this paper cites.
Optimization of computer simulation models with rare events
R. Y. Rubinstein · 1997
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
Learning to select and generalize striking movements in robot table tennis
K. Mülling, J. Kober, O. Kroemer, and J. Peters · 2013
Earlier work this paper cites.
Model-based hierarchical reinforcement learning and human action control
M. Botvinick and A. Weinstein · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. Theodorou · 2015
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
beta-VAE: Learning basic visual concepts with a constrained variational framework
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Taco: Learning task decomposition via temporal alignment for control
K. Shiarlis, M. Wulfmeier, S. Salter, S. Whiteson, and I. Posner · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Earlier work this paper cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. P. Lillicrap, and M. A. Riedmiller · 2018
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2019
Cited alongside, same era.
Latent skill planning for exploration and transfer
K. Xie, H. Bharadhwaj, D. Hafner, A. Garg, and F. Shkurti · 2020
Later among the works it cites.
Discovering motor programs by recomposing demonstrations
T. Shankar, S. Tulsiani, L. Pinto, and A. Gupta · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Reset-free lifelong learning with skill-space planning
K. Lu, A. Grover, P. Abbeel, and I. Mordatch · 2021
Later among the works it cites.
Model-based offline planning
A. Argenson and G. Dulac-Arnold · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Composing complex skills by learning transition policies
Y. Lee, S.-H. Sun, S. Somasundaram, E. S. Hu, and J. J. Lim · 2019
Cited alongside, same era.
Why does hierarchy (sometimes) work so well in reinforcement learning?
O. Nachum, H. Tang, X. Lu, S. Gu, H. Lee, and S. Levine · 2019
Cited alongside, same era.
Compile: Compositional imitation learning and execution
T. Kipf, Y. Li, H. Dai, V. Zambaldi, A. Sanchez-Gonzalez, E. Grefenstette, P. Kohli, and P. Battaglia · 2019
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Cited alongside, same era.
Learning to coordinate manipulation skills via skill behavior diversification
Y. Lee, J. Yang, and J. J. Lim · 2020
Cited alongside, same era.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Cited alongside, same era.
Later among the works it cites.
Adversarial skill chaining for long-horizon robot manipulation via terminal state regularization
Y. Lee, J. J. Lim, A. Anandkumar, and Y. Zhu · 2021
Later among the works it cites.
Accelerating robotic reinforcement learning via parameterized action primitives
M. Dalal, D. Pathak, and R. Salakhutdinov · 2021
Later among the works it cites.
Demonstration-guided reinforcement learning with learned skills
K. Pertsch, Y. Lee, Y. Wu, and J. J. Lim · 2021
Later among the works it cites.
Discovering and achieving goals via world models
R. Mendonca, O. Rybkin, K. Daniilidis, D. Hafner, and D. Pathak · 2021
Later among the works it cites.
Learning task decomposition with ordered memory policy network
Y. Lu, Y. Shen, S. Zhou, A. Courville, J. B. Tenenbaum, and C. Gan · 2021
Later among the works it cites.
Example-driven model-based reinforcement learning for solving long-horizon visuomotor tasks
B. Wu, S. Nair, L. Fei-Fei, and C. Finn · 2021
Later among the works it cites.
Temporal difference learning for model predictive control
N. Hansen, X. Wang, and H. Su · 2022
Closest in time.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
D. Shah, A. T. Toshev, S. Levine, and brian ichter · 2022
Closest in time.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Closest in time.