Fetching the paper…
Reading the bibliography…
Solving a reinforcement learning (RL) problem poses two competing challenges: fitting a potentially discontinuous value function, and generalizing well to new observations.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
Christopher J C H Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Advantage updating
L Baird · 1993
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Proto-value functions: Developmental reinforcement learning
Sridhar Mahadevan · 2005
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Graph-based regularization for spherical signal interpolation
Tamara Tošić and Pascal Frossard · 2010
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Policy distillation
Andrei A Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
A Laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
The hippocampus as a predictive map
Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Earlier work this paper cites.
Generalization and regularization in DQN
Jesse Farebrother, Marlos C Machado, and Michael Bowling · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Observe and look further: Achieving consistent performance on atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado Van Hasselt, John Quan, Mel Večerík, et al · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio · 2018
Cited alongside, same era.
Towards characterizing divergence in deep Q-learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Distilling policy distillation
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Generalization across space and time in reinforcement learning
Alex Lewandowski · 2020
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Later among the works it cites.
GRAC: self-guided and self-regularized actor-critic
Lin Shao, Yifan You, Mengyuan Yan, Qingyun Sun, and Jeannette Bohg · 2020
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech M Czarnecki, Razvan Pascanu, Simon Osindero, Siddhant Jayakumar, Grzegorz Swirszcz, and Max Jaderberg · 2019
Cited alongside, same era.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Cited alongside, same era.
SGD on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Cited alongside, same era.
Generalizing from a few environments in safety-critical reinforcement learning
Zachary Kenton, Angelos Filos, Owain Evans, and Yarin Gal · 2019
Cited alongside, same era.
Overcoming catastrophic interference in online reinforcement learning with dynamic self-organizing maps
Yat Long Lo and Sina Ghiassian · 2019
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle Pérez, Chico Q Camargo, and Ard A Louis · 2019
Cited alongside, same era.
Ray interference: a source of plateaus in deep reinforcement learning
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu · 2019
Cited alongside, same era.
Kaixin Wang, Bingyi Kang, Jie Shao, and Jiashi Feng · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2020
Later among the works it cites.
Invariant causal prediction for block MDPs
Amy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos, Marta Kwiatkowska, Joelle Pineau, Yarin Gal, and Doina Precup · 2020
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Later among the works it cites.
Fourier features in reinforcement learning with neural networks
David Brellmann, Goran Frehse, and David Filliat · 2021
Later among the works it cites.
Phasic policy gradient
Karl W Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2021
Later among the works it cites.
Batch normalization orthogonalizes representations in deep random networks
Hadi Daneshmand, Amir Joudaki, and Francis Bach · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation
Nicklas Hansen and Xiaolong Wang · 2021
Later among the works it cites.
Transient non-stationarity and generalisation in deep reinforcement learning
Maximilian Igl, Gregory Farquhar, Jelena Luketina, Wendelin Böhmer, and Shimon Whiteson · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
DR3: Value-based deep reinforcement learning requires explicit regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron Courville, George Tucker, and Sergey Levine · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney · 2021
Later among the works it cites.
The difficulty of passive learning in deep reinforcement learning
Georg Ostrovski, Pablo Samuel Castro, and Will Dabney · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Roberta Raileanu and Rob Fergus · 2021
Later among the works it cites.
Automatic data augmentation for generalization in reinforcement learning
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
Overcoming the spectral bias of neural value approximation
Ge Yang, Anurag Ajay, and Pulkit Agrawal · 2022
Closest in time.