Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (deep RL) has achieved tremendous success on various domains through a combination of algorithmic design and careful selection of hyper-parameters.
On inductive biases in deep reinforcement learning
Matteo Hessel, Hado van Hasselt, Joseph Modayil, and David Silver · 1907
Earlier work this paper cites.
A New Measure Of Rank Correlation
M. G. Kendall · 1938
Earlier work this paper cites.
The Problem of m m Rankings
M. G. Kendall and B. Babington Smith · 1939
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Python reference manual
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Michael J. Kearns and Satinder P. Singh · 2000
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
John D Hunter · 2007
Earlier work this paper cites.
Python for scientific computing
Travis E. Oliphant · 2007
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Python for Data Analysis: Data Wrangling with Pandas, NumPy, and IPython
Wes McKinney · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
Vincent François-Lavet, Raphael Fonteneau, and Damien Ernst · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Earlier work this paper cites.
Jupyter Notebooks—a publishing format for reproducible computational workflows
Thomas Kluyver, Benjain Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Jason Grout, Sylvain Corlay, Paul Ivanov, Damián Avila, Safia Abdalla, Carol Willing, and Jupyter Development Team · 2016
Earlier work this paper cites.
T. Schaul, John Quan, Ioannis Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and R. Munos · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George Driessche, Thore Graepel, and Demis Hassabis · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
Smooth activations and reproducibility in deep networks
G. Shamir, Dong Lin, and Lorenzo Coviello · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jax: composable transformations of python+ numpy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, et al · 2018
Cited alongside, same era.
Dopamine: A Research Framework for Deep Reinforcement Learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, E. Uchibe, and K. Doya · 2018
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alexander Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, H. V. Hasselt, T. Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, M. G. Azar, and D. Silver · 2018
Cited alongside, same era.
João Guilherme Madeira Araújo, Johan Samir Obando-Ceron, and Pablo Samuel Castro · 2021
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Later among the works it cites.
Evolving reinforcement learning algorithms
John D Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real, Sergey Levine, Quoc V Le, Honglak Lee, and Aleksandra Faust · 2021
Later among the works it cites.
Spectral normalisation for deep reinforcement learning: an optimisation perspective
Florin Gogianu, Tudor Berariu, Mihaela C Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Later among the works it cites.
Hyperparameter selection for imitation learning
Léonard Hussenot, Marcin Andrychowicz, Damien Vincent, Robert Dadashi, Anton Raichuk, Sabela Ramos, Nikola Momchev, Sertan Girgin, Raphael Marinier, Lukasz Stafiniak, et al · 2021
Later among the works it cites.
Revisiting design choices in offline model-based reinforcement learning
Cong Lu, Philip J Ball, Jack Parker-Holder, Michael A Osborne, and Stephen J Roberts · 2021
Later among the works it cites.
The difficulty of passive learning in deep reinforcement learning
Georg Ostrovski, Pablo Samuel Castro, and Will Dabney · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan D. Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, Craig Donner, Leslie Fritz, Cristian Galperti, Andrea Huber, James Keeling, Maria Tsimpoukelli, Jackie Kay, Antoine Merle, J-M. Moret, Seb Noury, Federico Pesamosca, David G. Pfau, Olivier Sauter, Cristian Sommariva, Stefano Coda, B. Duval, Ambrogio Fasoli, Pushmeet Kohli, Koray Kavukcuoglu, Demis Hassabis, and Martin A. Riedmiller · 2022
Later among the works it cites.
Proto-value networks: Scaling representation learning with auxiliary tasks
Jesse Farebrother, Joshua Greaves, Rishabh Agarwal, Charline Le Lan, Ross Goroshin, Pablo Samuel Castro, and Marc G Bellemare · 2022
Later among the works it cites.
The state of sparse training in deep reinforcement learning
Laura Graesser, Utku Evci, Erich Elsen, and Pablo Samuel Castro · 2022
Later among the works it cites.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Later among the works it cites.
Discovered policy optimisation
Chris Lu, Jakub Kuba, Alistair Letcher, Luke Metz, Christian Schroeder de Witt, and Jakob Foerster · 2022
Later among the works it cites.
Learning dynamics and generalization in deep reinforcement learning
Clare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska, and Yarin Gal · 2022
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Jack Parker-Holder, Raghu Rajan, Xingyou Song, André Biedenkapp, Yingjie Miao, Theresa Eimer, Baohe Zhang, Vu Nguyen, Roberto Calandra, Aleksandra Faust, et al · 2022
Later among the works it cites.
Investigating multi-task pretraining and generalization in reinforcement learning
Adrien Ali Taiga, Rishabh Agarwal, Jesse Farebrother, Aaron Courville, and Marc G Bellemare · 2022
Later among the works it cites.
Hyperparameters in reinforcement learning and how to tune them
Theresa Eimer, Marius Lindauer, and Roberta Raileanu · 2023
Later among the works it cites.
Small batch deep reinforcement learning
Johan Obando Ceron, Marc Bellemare, and Pablo Samuel Castro · 2023
Later among the works it cites.
Bigger, better, faster: Human-level Atari with human-level efficiency
Max Schwarzer, Johan Samir Obando-Ceron, Aaron Courville, Marc G Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, and Utku Evci · 2023
Later among the works it cites.
Stop regressing: Training value functions via classification for scalable deep rl
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, and Rishabh Agarwal · 2024
Closest in time.