Fetching the paper…
Reading the bibliography…
Experience replay enables off-policy reinforcement learning (RL) agents to utilize past experiences to maximize the cumulative reward.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P · 1992
Earlier work this paper cites.
A neural substrate of prediction and reward
Schultz, W., Dayan, P., and Montague, P. R · 1997
Earlier work this paper cites.
Understanding dopamine and reinforcement learning: the dopamine reward prediction error hypothesis
Glimcher, P. W · 2011
Earlier work this paper cites.
The role of rewarding and novel events in facilitating memory persistence in a separate spatial memory task
Salvetti, B., Morris, R. G., and Wang, S.-H · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Post-learning Hippocampal Dynamics Promote Preferential Retention of Rewarding Events
Gruber, M. J., Ritchey, M., Wang, S.-f., Doss, M. K., and Ranganath, C · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
Optimization models and applications
El Ghaoui, L · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Cited alongside, same era.
Pan, Y., Zaheer, M., White, A., Patterson, A., and White, M · 2018
Later among the works it cites.
Human hippocampal replay during rest prioritizes weakly learned information and predicts memory performance
Schapiro, A. C., McDevitt, E. A., Rogers, T. T., Mednick, S. C., and Norman, K. A · 2018
Later among the works it cites.
Reconciling λ \lambda -returns with experience replay
Daley, B. and Amato, C · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Later among the works it cites.
Post-learning hippocampal replay selectively reinforces spatial memory for highly rewarded locations
Michon, F., Sun, J.-J., Kim, C. Y., Ciliberti, D., and Kloosterman, F · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Cited alongside, same era.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Cited alongside, same era.
Prioritized memory access explains planning and hippocampal replay
Mattar, M. G. and Daw, N. D · 2018
Cited alongside, same era.
The role of hippocampal replay in memory and planning
Ólafsdóttir, H. F., Bush, D., and Barry, C · 2018
Cited alongside, same era.
Behavioural and computational evidence for memory consolidation biased by reward-prediction errors
Roscow, E. L., Jones, M. W., and Lepora, N. F · 2019
Later among the works it cites.
Importance resampling for off-policy prediction
Schlegel, M., Chung, W., Graves, D., Qian, J., and White, M · 2019
Later among the works it cites.
Experience replay optimization
Zha, D., Lai, K.-H., Zhou, K., and Hu, X · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.