Fetching the paper…
Reading the bibliography…
Experience replay (ER) is a fundamental component of off-policy deep reinforcement learning (RL).
A numerical method for solving incompressible viscous flow problems
A. J. Chorin · 1967
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L. H. Lin · 1992
Earlier work this paper cites.
A penalization method to take into account obstacles in incompressible viscous flows
P. Angot, C. Bruneau, and P. Fabrie · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
The matrix cookbook
K. B. Petersen, M. S. Pedersen, et al · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
T. Degris, M. White, and R. S. Sutton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
The importance of experience replay database composition in deep reinforcement learning
T. de Bruin, J. Kober, K. Tuyls, and R. Babuška · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, j. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Synchronisation through learning for two self-propelled swimmers
G. Novati, S. Verma, D. Alexeev, D. Rossinelli, W. M. van Rees, and P. Koumoutsakos · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
A. Rajeswaran, K. Lowrey, E. V. Todorov, and S. M. Kakade · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Koray, and N. de Freitas · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to soar in turbulent environments
G. Reddy, A. Celani, T. J. Sejnowski, and M. Vergassola · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Flow navigation by smart microswimmers via reinforcement learning
S. Colabrese, K. Gustavsson, A. Celani, and L. Biferale · 2017
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
S. Gu, T. Lillicrap, Z. Ghahramani, R. E Turner, and S. Levine · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Cited alongside, same era.
Closest in time.
Recall traces: Backtracking models for efficient reinforcement learning
A. Goyal, P. Brakel, W. Fedus, T. Lillicrap, S. Levine, H. Larochelle, and Y. Bengio · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Closest in time.
Selective experience replay for lifelong learning
D. Isele and A. Cosgun · 2018
Closest in time.
J. Oh, Y. Guo, S. Singh, and H. Lee · 2018
Closest in time.
Openai five
OpenAI · 2018
Closest in time.
Y. Pan, M. Zaheer, A. White, A. Patterson, and M. White · 2018
Closest in time.
The mirage of action-dependent baselines in reinforcement learning
G. Tucker, S. Bhupatiraju, S. Gu, R. E Turner, Z. Ghahramani, and S. Levine · 2018
Closest in time.
Efficient collective swimming by harnessing vortices through deep reinforcement learning
S. Verma, G. Novati, and P. Koumoutsakos · 2018
Closest in time.