Fetching the paper…
Reading the bibliography…
A major component of overfitting in model-free reinforcement learning (RL) involves the case where the agent may mistakenly correlate reward with certain spurious features from the observations generated by the Markov Decision Process (MDP).
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 1902
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 1905
Earlier work this paper cites.
Weight agnostic neural networks
Adam Gaier and David Ha · 1906
Earlier work this paper cites.
Martín Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 1907
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir N Vapnik and A Ya Chervonenkis · 1971
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David A. McAllester · 2003
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M. Kakade, Karthik Sridharan, and Ambuj Tewari · 2008
Earlier work this paper cites.
Lectures in geometric functional analysis
Roman Vershynin · 2009
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
Joseph K. Bradley, Aapo Kyrola, Danny Bickson, and Carlos Guestrin · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel Todorov, and Sham M. Kakade · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Starcraft II: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, John Quan, Stephen Gaffney, Stig Petersen, Karen Simonyan, Tom Schaul, Hado van Hasselt, David Silver, Timothy P. Lillicrap, Kevin Calderone, Paul Keet, Anthony Brunasso, David Lawrence, Anders Ekermo, Jacob Repp, and Rodney Tsing · 2017
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Later among the works it cites.
Gotta learn fast: A new benchmark for generalization in RL
Alex Nichol, Vicki Pfau, Christopher Hesse, Oleg Klimov, and John Schulman · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C. Karen Liu, and Greg Turk · 2017
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Cited alongside, same era.
Lipschitz continuity in model-based reinforcement learning
Kavosh Asadi, Dipendra Misra, and Michael L. Littman · 2018
Cited alongside, same era.
Structured evolution with compact architectures for scalable policy optimization
Krzysztof Choromanski, Mark Rowland, Vikas Sindhwani, Richard E. Turner, and Adrian Weller · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Later among the works it cites.
Relational recurrent neural networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack W. Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, and Timothy P. Lillicrap · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le · 2018
Later among the works it cites.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang, Pablo Samuel Castro, Yulun Li, Ludwig Schubert, Marc G. Bellemare, Jeff Clune, and Joel Lehman · 2018
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2019
Closest in time.
Transfer learning for related reinforcement learning tasks via image-to-image translation
Shani Gamrian and Yoav Goldberg · 2019
Closest in time.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2019
Closest in time.
Conditional variance penalties and domain shift robustness
Christina Heinze-Deml and Nicolai Meinshausen · 2019
Closest in time.
Invariant causal prediction for nonlinear models
Christina Heinze-Deml, Jonas Peters, and Nicolai Meinshausen · 2019
Closest in time.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Closest in time.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso A. Poggio, Alexander Rakhlin, and James Stokes · 2019
Closest in time.
Generalization in deep networks: The role of distance from initialization
Vaishnavh Nagarajan and J Zico Kolter · 2019
Closest in time.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Stephen Tu and Benjamin Recht · 2019
Closest in time.
On the generalization gap in reparameterizable reinforcement learning
Huan Wang, Stephan Zheng, Caiming Xiong, and Richard Socher · 2019
Closest in time.