Fetching the paper…
Reading the bibliography…
Plasticity refers to a network's ability to adapt to changing data distributions, which is crucial for the successful training of deep reinforcement learning agents.
Solving Rubik’s Cube with a Robot Hand
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang. 2019 · 1910
Earlier work this paper cites.
Dota 2 with Large Scale Deep Reinforcement Learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Christopher Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang. 2019 · 1912
Earlier work this paper cites.
Reinforcement learning with selective perception and hidden state
Andrew Kachites McCallum. 1996 · 1996
Earlier work this paper cites.
Flat Minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions
Sepp Hochreiter. 1998 · 1998
Earlier work this paper cites.
Reinforcement learning - an introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. 2020 · 2005
Earlier work this paper cites.
Double Q-learning. In Advances in Neural Information Processing Systems (NeurIPS) , John D. Lafferty, Christopher K. I. Williams, John Shawe-Taylor, Richard S. Zemel, and Aron Culotta (Eds.). Curran Associates, Inc., 2613–2621
Hado van Hasselt. 2010 · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2012, Vilamoura, Algarve, Portugal, October 7-12, 2012 . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models. In International Conference on Machine Learning (ICML) (JMLR Workshop and Conference Proceedings, Vol. 28) . JMLR.org
Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al · 2013
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In International Conference on Machine Learning (ICML) (JMLR Workshop and Conference Proceedings, Vol. 37) , Francis R. Bach and David M. Blei (Eds.). JMLR.org, 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units. In International Conference on Machine Learning (ICML) (JMLR Workshop and Conference Proceedings, Vol. 48) , Maria-Florina Balcan and Kilian Q. Weinberger (Eds.). JMLR.org, 2217–2225
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
A Distributional Perspective on Reinforcement Learning. In International Conference on Machine Learning (ICML) , Vol. 70. 449–458
Marc G. Bellemare, Will Dabney, and Rémi Munos. 2017 · 2017
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples. In International conference on machine learning . PMLR, 854–863
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. 2017 · 2017
Earlier work this paper cites.
Count-Based Exploration with Neural Density Models. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 70) , Doina Precup and Yee Whye Teh (Eds.). PMLR, 2721–2730
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 6306–6315
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures. In International Conference on Machine Learning (ICML) . 1406–1415
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu. 2018 · 2018
Earlier work this paper cites.
Addressing Function Approximation Error in Actor-Critic Methods. In International Conference on Machine Learning (ICML) . 1582–1591
Scott Fujimoto, Herke van Hoof, and David Meger. 2018 · 2018
Earlier work this paper cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In International Conference on Machine Learning (ICML) . 1856–1865
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Improving Regression Performance with Distributional Losses. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 80) , Jennifer G. Dy and Andreas Krause (Eds.). PMLR, 2162–2171
Ehsan Imani and Martha White. 2018 · 2018
Earlier work this paper cites.
Spectral Normalization for Generative Adversarial Networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018 · 2018
Earlier work this paper cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller. 2018 · 2018
Earlier work this paper cites.
Deep Reinforcement Learning and the Deadly Triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil. 2018 · 2018
Earlier work this paper cites.
Continual Learning with Neural Networks: A Review. In Proceedings of the ACM India Joint International Conference on Data Science and Management of Data, COMAD/CODS 2019, Kolkata, India, January 3-5, 2019 , Raghu Krishnapuram and Parag Singla (Eds.). ACM, 362–365
Abhijeet Awasthi and Sunita Sarawagi. 2019 · 2019
Earlier work this paper cites.
An Evaluation of Parametric Activation Functions for Deep Learning. In International Conference on Systems, Man and Cybernetics (SMC) . IEEE, 3006–3011
Luke B. Godfrey. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32 (NeurIPS) , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 8024–8035
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Understanding and Improving Layer Normalization. In Advances in Neural Information Processing Systems 32 (NeurIPS) , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 4383–4393
Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin. 2019 · 2019
Earlier work this paper cites.
On Warm-Starting Neural Network Training. In Advances in Neural Information Processing Systems (NeurIPS)
Jordan T. Ash and Ryan P. Adams. 2020 · 2020
Earlier work this paper cites.
Model Based Reinforcement Learning for Atari. In International Conference on Learning Representations (ICLR) . OpenReview.net
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H. Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, and Henryk Michalewski. 2020 · 2020
Earlier work this paper cites.
Reinforcement Learning with Augmented Data. In Advances in Neural Information Processing Systems (NeurIPS) , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas. 2020 · 2020
Earlier work this paper cites.
Mastering Atari, Go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy P. Lillicrap, and David Silver. 2020 · 2020
Cited alongside, same era.
What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale Study. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net
Marcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, and Olivier Bachem. 2021 · 2021
Cited alongside, same era.
A study on the plasticity of neural networks
Tudor Berariu, Wojciech Czarnecki, Soham De, Jörg Bornschein, Samuel L. Smith, Razvan Pascanu, and Claudia Clopath. 2021 · 2021
Cited alongside, same era.
Towards Deeper Deep Reinforcement Learning with Spectral Normalization
Johan Bjorck, Carla P. Gomes, and Kilian Q. Weinberger. 2021 · 2021
Cited alongside, same era.
Small batch deep reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS) , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Johan S. Obando-Ceron, Marc G. Bellemare, and Pablo Samuel Castro. 2023 · 2023
Later among the works it cites.
Bigger, Better, Faster: Human-level Atari with human-level efficiency. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 30365–30380
Max Schwarzer, Johan Samir Obando-Ceron, Aaron C. Courville, Marc G. Bellemare, Rishabh Agarwal, and Pablo Samuel Castro. 2023 · 2023
Later among the works it cites.
The Dormant Neuron Phenomenon in Deep Reinforcement Learning. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 32145–32168
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, and Utku Evci. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Randomized Ensembled Double Q-Learning: Learning Fast Without a Model. In International Conference on Learning Representations (ICLR) . OpenReview.net
Xinyue Chen, Che Wang, Zijian Zhou, and Keith W. Ross. 2021 · 2021
Cited alongside, same era.
Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations (ICLR)
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021 · 2021
Cited alongside, same era.
Spectral Normalisation for Deep Reinforcement Learning: An Optimisation Perspective. In International Conference on Machine Learning (ICML) . 3734–3744
Florin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu. 2021 · 2021
Cited alongside, same era.
Transient Non-stationarity and Generalisation in Deep Reinforcement Learning. In International Conference on Learning Representations (ICLR) . OpenReview.net
Maximilian Igl, Gregory Farquhar, Jelena Luketina, Wendelin Boehmer, and Shimon Whiteson. 2021 · 2021
Cited alongside, same era.
Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning. In International Conference on Learning Representations (ICLR) . OpenReview.net
Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine. 2021 · 2021
Cited alongside, same era.
On the Effect of Auxiliary Tasks on Representation Dynamics. In International Conference on Artificial Intelligence and Statistics (AISTATS) (Proceedings of Machine Learning Research, Vol. 130) , Arindam Banerjee and Kenji Fukumizu (Eds.). PMLR, 1–9
Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney. 2021 · 2021
Cited alongside, same era.
The Difficulty of Passive Learning in Deep Reinforcement Learning. In Advances in Neural Information Processing Systems 34 (NeurIPS) , Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 23283–23295
Georg Ostrovski, Pablo Samuel Castro, and Will Dabney. 2021 · 2021
Cited alongside, same era.
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels. In International Conference on Learning Representations (ICLR) . OpenReview.net
Denis Yarats, Ilya Kostrikov, and Rob Fergus. 2021 · 2021
Cited alongside, same era.
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity. In International Conference on Learning Representations (ICLR)
Aditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, and Jan Peters. 2024 · 2024
Closest in time.
Parseval Regularization for Continual Reinforcement Learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Wesley Chung, Lynn Cherif, Doina Precup, and David Meger. 2024 · 2024
Closest in time.
Adaptive Rational Activations to Boost Deep Reinforcement Learning. In International Conference on Learning Representations (ICLR) . OpenReview.net
Quentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina, and Kristian Kersting. 2024 · 2024
Closest in time.
Loss of plasticity in deep continual learning
Shibhansh Dohare, J. Fernando Hernandez-Garcia, Qingfeng Lan, Parash Rahman, A. Rupam Mahmood, and Richard S. Sutton. 2024 · 2024
Closest in time.
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Benjamin Ellis, Matthew Thomas Jackson, Andrei Lupu, Alexander David Goldie, Mattie Fellows, Shimon Whiteson, and Jakob N. Foerster. 2024 · 2024
Closest in time.
Weight Clipping for Deep Continual and Reinforcement Learning
Mohamed Elsayed, Qingfeng Lan, Clare Lyle, and A. Rupam Mahmood. 2024a · 2024
Closest in time.
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning. In International Conference on Learning Representations (ICLR)
Mohamed Elsayed and A. Rupam Mahmood. 2024 · 2024
Closest in time.
Streaming Deep Reinforcement Learning Finally Works
Mohamed Elsayed, Gautham Vasan, and A. Rupam Mahmood. 2024b · 2024
Closest in time.
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL. In International Conference on Machine Learning (ICML)
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, and Rishabh Agarwal. 2024 · 2024
Closest in time.
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Alexandre Galashov, Michalis K. Titsias, András György, Clare Lyle, Razvan Pascanu, Yee Whye Teh, and Maneesh Sahani. 2024 · 2024
Closest in time.
Can Learned Optimization Make Reinforcement Learning Less Difficult?. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson, and Jakob N. Foerster. 2024 · 2024
Closest in time.
Adaptive Regularization of Representation Rank as an Implicit Constraint of Bellman Equation. In International Conference on Learning Representations (ICLR)
Qiang He, Tianyi Zhou, Meng Fang, and Setareh Maghsudi. 2024 · 2024
Closest in time.
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
Marcel Hussing, Claas Voelcker, Igor Gilitschenski, Amir-massoud Farahmand, and Eric Eaton. 2024 · 2024
Closest in time.
ACE: Off-Policy Actor-Critic with Causality-Aware Entropy Regularization. In International Conference on Machine Learning (ICML) . OpenReview.net
Tianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo, Guowei Xu, Jiawei Guo, Ruijie Zheng, Furong Huang, Fuchun Sun, and Huazhe Xu. 2024 · 2024
Closest in time.
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Arthur Juliani and Jordan T. Ash. 2024 · 2024
Closest in time.
Hadamard Representations: Augmenting Hyperbolic Tangents in RL
Jacob E Kooi, Mark Hoogendoorn, and Vincent François-Lavet. 2024 · 2024
Closest in time.
Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks. In International Conference on Machine Learning (ICML) . OpenReview.net
Hojoon Lee, Hyeonseo Cho, Hyunseung Kim, Donghu Kim, Dugki Min, Jaegul Choo, and Clare Lyle. 2024 · 2024
Closest in time.
Normalization and effective learning rates in reinforcement learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens, Hado Philip van Hasselt, Razvan Pascanu, and Will Dabney. 2024a · 2024
Closest in time.
Disentangling the Causes of Plasticity Loss in Neural Networks. In Conference on Lifelong Learning Agents, 29-1 August 2024, University of Pisa, Pisa, Italy (Proceedings of Machine Learning Research, Vol. 274) , Vincenzo Lomonaco, Stefano Melacci, Tinne Tuytelaars, Sarath Chandar, and Razvan Pascanu (Eds.). PMLR, 750–783
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, Hado van Hasselt, Razvan Pascanu, James Martens, and Will Dabney. 2024b · 2024
Closest in time.
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages. In International Conference on Learning Representations (ICLR) . OpenReview.net
Guozheng Ma, Lu Li, Sen Zhang, Zixuan Liu, Zhen Wang, Yixin Chen, Li Shen, Xueqian Wang, and Dacheng Tao. 2024 · 2024
Closest in time.
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Skander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu, and Caglar Gulcehre. 2024 · 2024
Closest in time.
Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos, and Marek Cygan. 2024b · 2024
Closest in time.
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
Johan S. Obando-Ceron, João Guilherme Madeira Araújo, Aaron C. Courville, and Pablo Samuel Castro. 2024a · 2024
Closest in time.
Mixtures of Experts Unlock Parameter Scaling for Deep RL. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net
Johan Samir Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle, Jesse Farebrother, Jakob Nicolaus Foerster, Gintare Karolina Dziugaite, Doina Precup, and Pablo Samuel Castro. 2024c · 2024
Closest in time.
Continual Learning of Large Language Models: A Comprehensive Survey
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, and Hao Wang. 2024 · 2024
Closest in time.
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.)
Gautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He, Fahim Shahriar, Colin Bellinger, Martha White, and Rupam Mahmood. 2024 · 2024
Closest in time.
A Comprehensive Survey of Continual Learning: Theory, Method and Application
Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. 2024 · 2024
Closest in time.
DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization. In International Conference on Learning Representations (ICLR) . OpenReview.net
Guowei Xu, Ruijie Zheng, Yongyuan Liang, Xiyao Wang, Zhecheng Yuan, Tianying Ji, Yu Luo, Xiaoyu Liu, Jiaxin Yuan, Pu Hua, Shuzhen Li, Yanjie Ze, Hal Daumé III, Furong Huang, and Huazhe Xu. 2024 · 2024
Closest in time.
Simplifying Deep Temporal Difference Learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Matteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou, Ivan Masmitja, Jakob Nicolaus Foerster, and Mario Martin. 2025a · 2025
Closest in time.
Simplifying Deep Temporal Difference Learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Matteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou, Ivan Masmitja, Jakob Nicolaus Foerster, and Mario Martin. 2025b · 2025
Closest in time.
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Hojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone, and Takuma Seno. 2025 · 2025
Closest in time.
Learning Continually by Spectral Regularization. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Alex Lewandowski, Michal Bortkiewicz, Saurabh Kumar, András György, Dale Schuurmans, Mateusz Ostaszewski, and Marlos C. Machado. 2025a · 2025
Closest in time.
Plastic Learning with Deep Fourier Features. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Alex Lewandowski, Dale Schuurmans, and Marlos C. Machado. 2025b · 2025
Closest in time.
Neuroplastic Expansion in Deep Reinforcement Learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Jiashun Liu, Johan S. Obando-Ceron, Aaron C. Courville, and Ling Pan. 2025 · 2025
Closest in time.
Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement Learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net
Théo Vincent, Fabian Wahren, Jan Peters, Boris Belousov, and Carlo D’Eramo. 2025 · 2025
Closest in time.
Leveraging Procedural Generation to Benchmark Reinforcement Learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 2048–2056
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman. 2020 · 2056
Closest in time.