Fetching the paper…
Reading the bibliography…
Several algorithms have been proposed to sample non-uniformly the replay buffer of deep Reinforcement Learning (RL) agents to speed-up learning, but very few theoretical foundations of these sampling schemes have been provided.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W. and Atkeson, C. G · 1993
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R · 2005
Earlier work this paper cites.
Approximate modified policy iteration
Scherrer, B., Ghavamzadeh, M., Gabillon, V., and Geist, M · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Needell, D., Ward, R., and Srebro, N · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, P. and Zhang, T · 2015
Earlier work this paper cites.
Variance reduction in sgd by distributed importance sampling
Alain, G., Lamb, A., Sankar, C., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Online batch selection for faster training of neural networks
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Simulation and the Monte Carlo method
Rubinstein, R. Y. and Kroese, D. P · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
PyBullet, a Python module for physics simulation for games, robotics and machine learning
Coumans, E. and Bai, Y · 2019
Later among the works it cites.
State distribution-aware sampling for deep q-learning
Li, W., Huang, F., Li, X., Pan, G., and Wu, F · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerating deep neural network training with inconsistent stochastic gradient descent
Wang, L., Yang, Y., Min, R., and Chakradhar, S · 2017
Cited alongside, same era.
A deeper look at experience replay
Zhang, S. and Sutton, R. S · 2017
Cited alongside, same era.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
The reactor: A fast and sample-efficient actor-critic agent for reinforcement learning
Gruslys, A., Dabney, W., Azar, M. G., Piot, B., Bellemare, M., and Munos, R · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Later among the works it cites.
Stable baselines3
Raffin, A., Hill, A., Ernestus, M., Gleave, A., Kanervisto, A., and Dormann, N · 2019
Later among the works it cites.
Boosting soft actor-critic: Emphasizing recent experience without forgetting the past
Wang, C. and Ross, K · 2019
Later among the works it cites.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Young, K. and Tian, T · 2019
Later among the works it cites.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W · 2020
Later among the works it cites.
An equivalence between loss functions and non-uniform sampling in experience replay
Fujimoto, S., Meger, D., and Precup, D · 2020
Later among the works it cites.
Discor: Corrective feedback in reinforcement learning via distribution correction
Kumar, A., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Obando-Ceron, J. S. and Castro, P. S · 2021
Closest in time.