Fetching the paper…
Reading the bibliography…
Off-policy deep reinforcement learning (RL) has been successful in a range of challenging domains.
The jackknife, the bootstrap, and other resampling plans , volume 38
Efron, B · 1982
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Ensemble algorithms in reinforcement learning
Wiering, M. A. and Van Hasselt, H · 2008
Earlier work this paper cites.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y., Munos, R., and Szepesvári, C · 2009
Earlier work this paper cites.
Double q-learning
Hasselt, H. V · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Later among the works it cites.
Revisiting the softmax bellman operator: New benefits and new perspective
Song, Z., Parr, R., and Carin, L · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Wang, T., Bao, X., Clavera, I., Hoang, J., Wen, Y., Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Addressing function approximation error in actor-critic methods
Fujimoto, S., Van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Universal planning networks
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C · 2018
Cited alongside, same era.
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Closest in time.
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B., Stanton, S., Yarats, D., and Wilson, A. G · 2020
Closest in time.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Closest in time.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2020
Closest in time.
Discor: Corrective feedback in reinforcement learning via distribution correction
Kumar, A., Gupta, A., and Levine, S · 2020
Closest in time.
Maxmin q-learning: Controlling the estimation bias of q-learning
Lan, Q., Pan, Y., Fyshe, A., and White, M · 2020
Closest in time.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Closest in time.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X., Nagabandi, A., Abbeel, P., and Levine, S · 2020
Closest in time.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, A., Laskin, M., and Abbeel, P · 2020
Closest in time.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2020
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Closest in time.