Fetching the paper…
Reading the bibliography…
Episodic memory-based methods can rapidly latch onto past successful strategies by a non-parametric memory and improve sample efficiency of traditional reinforcement learning.
Dynamic programming princeton university press princeton
Bellman, R · 1957
Earlier work this paper cites.
The art and thought of Heraclitus: a new arrangement and translation of the Fragments with literary and philosophical commentary
Kahn, C. H. et al · 1981
Earlier work this paper cites.
Configural association theory: The role of the hippocampal formation in learning, memory, and amnesia
Sutherland, R. J. and Rudy, J. W · 1989
Earlier work this paper cites.
Simple memory: a theory for archicortex
Marr, D., Willshaw, D., and McNaughton, B · 1991
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 1993
Earlier work this paper cites.
Case-based decision theory
Gilboa, I. and Schmeidler, D · 1995
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, S. P. and Sutton, R. S · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., Mansour, Y., et al · 1999
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś · 2002
Earlier work this paper cites.
Hippocampal contributions to control: The third way
Lengyel, M. and Dayan, P · 2007
Earlier work this paper cites.
Integrating memories in the human brain: hippocampal-midbrain encoding of overlapping events
Shohamy, D. and Wagner, A. D · 2008
Earlier work this paper cites.
Policy gradient methods
Peters, J. and Bagnell, J. A · 2010
Earlier work this paper cites.
Double q-learning
van Hasselt, H · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. A · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Blundell, C., Uria, B., Pritzel, A., Li, Y., Ruderman, A., Leibo, J. Z., Rae, J., Wierstra, D., and Hassabis, D · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Fast deep reinforcement learning using online adjustments from the past
Hansen, S., Pritzel, A., Sprechmann, P., Barreto, A., and Blundell, C · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Later among the works it cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Later among the works it cites.
Episodic memory deep q-networks
Lin, Z., Zhao, T., Yang, G., and Zhang, L · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Cited alongside, same era.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J · 2017
Cited alongside, same era.
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Convergent TREE BACKUP and RETRACE with function approximation
Touati, A., Bacon, P., Precup, D., and Vincent, P · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Later among the works it cites.
Reinforcement learning, fast and slow
Botvinick, M., Ritter, S., Wang, J. X., Kurth-Nelson, Z., Blundell, C., and Hassabis, D · 2019
Later among the works it cites.
Sample-efficient deep reinforcement learning via episodic backward update
Lee, S. Y., Choi, S., and Chung, S · 2019
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Kuznetsov, A., Shvechikov, P., Grishin, A., and Vetrov, D. P · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Lan, Q., Pan, Y., Fyshe, A., and White, M · 2020
Later among the works it cites.
Episodic reinforcement learning with associative memory
Zhu, G., Lin, Z., Yang, G., and Zhang, C · 2020
Later among the works it cites.
Randomized ensembled double q-learning: Learning fast without a model
Chen, X., Wang, C., Zhou, Z., and Ross, K · 2021
Closest in time.