Soft actor-critic algorithms and applications
Original
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Are deep policy gradient algorithms truly policy gradient algorithms?
Original
A. Ilyas, L. Engstrom, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. J. Storkey, T. Darrell, and A. A. Efros · 2019
Later among the works it cites.
If maxent rl is the answer, what is the question?, 2019
B. Eysenbach and S. Levine · 2019
Later among the works it cites.
Learning to predict without looking ahead: World models without forward prediction
D. Freeman, D. Ha, and L. Metz · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Later among the works it cites.
Explicit explore-exploit algorithms in continuous state spaces
M. Henaff · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Original
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation, 2019
A. Nagabandi, K. Konoglie, S. Levine, and V. Kumar · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 2019
Later among the works it cites.
Active inference: demystified and compared
Original
N. Sajid, P. J. Ball, and K. J. Friston · 2019
Later among the works it cites.
Mastering Atari, Go, chess and shogi by planning with a learned model, 2019
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2019
Later among the works it cites.
Model-based active exploration
P. Shyam, W. Jaskowski, and F. Gomez · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Original
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Closest in time.
Model based reinforcement learning for atari
Ł. Kaiser, M. Babaeizadeh, P. Miłos, B. Osiński, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker, and H. Michalewski · 2020
Closest in time.