Fetching the paper…
Reading the bibliography…
Learning efficiently from small amounts of data has long been the focus of model-based reinforcement learning, both for the online case when interacting with the environment and the offline case when learning from a fixed dataset.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
RL Unplugged: Benchmarks for Offline Reinforcement Learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T. L., Colmenarejo, S. G., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., Dulac-Arnold, G., Li, J., Norouzi, M., Hoffman, M., Nachum, O., Tucker, G., Heess, N., and de Freitas, N · 2006
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Layer normalization, 2016
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Machado, M., Bellemare, M., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
Distributed Distributional Deterministic Policy Gradients, 2018
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2018
Cited alongside, same era.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Haiku: Sonnet for JAX, 2020
Hennigan, T., Cai, T., Norman, T., and Babuschkin, I · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
MOReL : Model-Based Offline Reinforcement Learning, 2020
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Later among the works it cites.
Conservative Q-learning for Offline Reinforcement Learning, 2020
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization, 2020
Matsushima, T., Furuta, H., Matsuo, Y., Nachum, O., and Gu, S · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Deepmind control suite, 2018
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Cited alongside, same era.
Off-policy actor-critic with shared experience replay
Schmitt, S., Hessel, M., and Simonyan, K · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Later among the works it cites.
Offline reinforcement learning from images with latent space models
Rafailov, R., Yu, T., Rajeswaran, A., and Finn, C · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T. P., and Silver, D · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., and Riedmiller, M · 2020
Later among the works it cites.
Critic Regularized Regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al · 2020
Later among the works it cites.
MOPO: Model-based Offline Policy Optimization, 2020
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
POPO: Pessimistic Offline Policy Optimization, 2021
He, Q. and Hou, X · 2021
Closest in time.
Learning and Planning in Complex Action Spaces
Hubert, T., Schrittwieser, J., Antonoglou, I., Barekatain, M., Schmitt, S., and Silver, D · 2021
Closest in time.