Fetching the paper…
Reading the bibliography…
We show that deep reinforcement learning algorithms can retain their ability to learn without resetting network parameters in settings where the number of gradient updates greatly exceeds the number of environment samples by combatting value function divergence.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Learning internal representations by error propagation , pp. 318–362
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A. Pomerleau · 1988
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
Q-learning with hidden-unit restarting
Charles Anderson · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Robot learning from demonstration
Christopher G. Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Estimator variance in reinforcement learning: Theoretical problems and practical solutions
Mark Pendrith and Malcolm Ryan · 1997
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Christopher M. Bishop · 2006
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Shie Mannor, Duncan Simester, Peng Sun, and John N. Tsitsiklis · 2007
Earlier work this paper cites.
Double q-learning
Hado van Hasselt · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Bias-corrected q-learning to control max-operator bias in q-learning
Donghun Lee, Boris Defourny, and Warren Buckler Powell · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adaptive step-sizes for reinforcement learning
William C Dabney · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Averaged-DQN: Variance reduction and stabilization for deep reinforcement learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Weighted double q-learning
Zongzhang Zhang, Zhiyuan Pan, and Mykel J. Kochenderfer · 2017
Cited alongside, same era.
Generalization and regularization in DQN
Jesse Farebrother, Marlos C. Machado, and Michael Bowling · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Later among the works it cites.
Regularization matters in policy optimization - an empirical study on continuous control
Zhuang Liu, Xuanlin Li, Bingyi Kang, and Trevor Darrell · 2021
Later among the works it cites.
Understanding and preventing capacity loss in reinforcement learning
Clare Lyle, Mark Rowland, and Will Dabney · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning
Ted Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel, and Michael Jordan · 2021
Later among the works it cites.
Ensemble bootstrapping for q-learning
Oren Peer, Chen Tessler, Nadav Merlis, and Ron Meir · 2021
Later among the works it cites.
Estimation error correction in deep reinforcement learning for deterministic actor-critic methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Reducing estimation bias via triplet-average deep deterministic policy gradient
Dongming Wu, Xingping Dong, Jianbing Shen, and Steven C. H. Hoi · 2019
Cited alongside, same era.
Understanding and improving layer normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin · 2019
Cited alongside, same era.
Baturay Saglam, Enes Duran, Dogan C. Cicek, Furkan B. Mutlu, and Suleyman S. Kozat · 2021
Later among the works it cites.
Is high variance unavoidable in RL? a case study in continuous control
Johan Bjorck, Carla P Gomes, and Kilian Q Weinberger · 2022
Later among the works it cites.
Dropout q-functions for doubly efficient reinforcement learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
The phenomenon of policy churn
Tom Schaul, Andre Barreto, John Quan, and Georg Ostrovski · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning, 2023
Zaheer Abbas, Rosie Zhao, Joseph Modayil, Adam White, and Marlos C. Machado · 2023
Later among the works it cites.
Resetting the optimizer in deep RL: An empirical study
Kavosh Asadi, Rasool Fakoor, and Shoham Sabach · 2023
Later among the works it cites.
Efficient online reinforcement learning with offline data
Philip J. Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine · 2023
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
Pierluca D’Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G Bellemare, and Aaron Courville · 2023
Later among the works it cites.
Sample-efficient and safe deep reinforcement learning via reset deep ensemble agents
Woojun Kim, Yongjae Shin, Jongeui Park, and Youngchul Sung · 2023
Later among the works it cites.
Bridging RL theory and practice with the effective horizon
Cassidy Laidlaw, Stuart Russell, and Anca Dragan · 2023
Later among the works it cites.
Efficient deep reinforcement learning requires regulating overfitting
Qiyang Li, Aviral Kumar, Ilya Kostrikov, and Sergey Levine · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Later among the works it cites.
Bigger, better, faster: Human-level Atari with human-level efficiency
Max Schwarzer, Johan Samir Obando Ceron, Aaron Courville, Marc G Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, and Utku Evci · 2023
Later among the works it cites.
Revisiting the minimalist approach to offline reinforcement learning
Denis Tarasov, Vladislav Kurenkov, Alexander Nikulin, and Sergey Kolesnikov · 2023
Later among the works it cites.
Cross q q : Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Aditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, and Jan Peters · 2024
Closest in time.
TD-MPC2: Scalable, robust world models for continuous control
Nicklas Hansen, Hao Su, and Xiaolong Wang · 2024
Closest in time.
Disentangling the causes of plasticity loss in neural networks, 2024
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, Hado van Hasselt, Razvan Pascanu, James Martens, and Will Dabney · 2024
Closest in time.
Overestimation, overfitting, and plasticity in actor-critic: the bitter lesson of reinforcement learning
Michal Nauman, Michał Bortkiewicz, Piotr Miłoś, Tomasz Trzcinski, Mateusz Ostaszewski, and Marek Cygan · 2024
Closest in time.
Deep reinforcement learning with plasticity injection
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, and André Barreto · 2024
Closest in time.