Fetching the paper…
Reading the bibliography…
Most reinforcement learning algorithms take advantage of an experience replay buffer to repeatedly train on samples the agent has observed in the past.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Longxin Lin · 2004
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Daniel Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2017
Earlier work this paper cites.
A novel ddpg method with prioritized experience replay
Yuenan Hou, Lifeng Liu, Qing Wei, Xudong Xu, and Chunlin Chen · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare · 2021
Later among the works it cites.
When does loss-based prioritization fail?
Niel Teng Hu, Xinyu Hu, Rosanne Liu, Sara Hooker, and Jason Yosinski · 2021
Later among the works it cites.
Large batch experience replay
Thibault Lahire, M. Geist, and E. Rachelson · 2021
Later among the works it cites.
Regret minimization experience replay in off-policy reinforcement learning
Xu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang, Feng Xu, and Yang Yu · 2021
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando-Ceron and Pablo Samuel Castro · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepmind control suite
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Cited alongside, same era.
Boosting soft actor-critic: Emphasizing recent experience without forgetting the past
Che Wang and Keith Ross · 2019
Cited alongside, same era.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Kenny Young and Tian Tian · 2019
Cited alongside, same era.
An equivalence between loss functions and non-uniform sampling in experience replay
Scott Fujimoto, David Meger, and Doina Precup · 2020
Cited alongside, same era.
Discor: Corrective feedback in reinforcement learning via distribution correction
Aviral Kumar, Abhishek Gupta, and Sergey Levine · 2020
Cited alongside, same era.
Experience replay with likelihood-free importance weights
Samarth Sinha, Jiaming Song, Animesh Garg, and S. Ermon · 2020
Cited alongside, same era.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Cited alongside, same era.
Later among the works it cites.
Topological experience replay
Zhang-Wei Hong, Tao Chen, Yen-Chen Lin, J. Pajarinen, and Pulkit Agrawal · 2022
Closest in time.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Sören Mindermann, Jan M Brauner, Muhammed T Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N Gomez, Adrien Morisot, Sebastian Farquhar, and Yarin Gal · 2022
Closest in time.
Model-augmented prioritized experience replay
Youngmin Oh, Jinwoo Shin, Eunho Yang, and Sung Ju Hwang · 2022
Closest in time.
The phenomenon of policy churn, 2022
Tom Schaul, André Barreto, John Quan, and Georg Ostrovski · 2022
Closest in time.
Efficient deep reinforcement learning requires regulating overfitting, 2023
Qiyang Li, Aviral Kumar, Ilya Kostrikov, and Sergey Levine · 2023
Closest in time.