Fetching the paper…
Reading the bibliography…
The ability for policies to generalize to new environments is key to the broad application of RL agents.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Ohad Shamir, Sivan Sabato, and Naftali Tishby · 2010
Earlier work this paper cites.
Learning to grasp under uncertainty
Freek Stulp, Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Shimon Whiteson, Brian Tanner, Matthew E Taylor, and Peter Stone · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
The transferability approach: Crossing the reality gap in evolutionary robotics
Sylvain Koos, Jean-Baptiste Mouret, and Stéphane Doncieux · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Transfer deep reinforcement learning in 3d environments: An empirical study
Devendra Singh Chaplot, Guillaume Lample, Kanthashree Mysore Sathyendra, and Ruslan Salakhutdinov · 2016
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matthew Botvinick · 2016
Cited alongside, same era.
Deep variational information bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy · 2017
Cited alongside, same era.
Reinforcement learning for pivoting task
Rika Antonova, Silvia Cruciani, Christian Smith, and Danica Kragic · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Illuminating generalization in deep reinforcement learning through procedural level generation
Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager, Ahmed Khalifa, Julian Togelius, and Sebastian Risi · 2018
Later among the works it cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Later among the works it cites.
Gotta learn fast: A new benchmark for generalization in rl
Alex Nichol, Vicki Pfau, Christopher Hesse, Oleg Klimov, and John Schulman · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George · 2017
Cited alongside, same era.
Adversarially robust policy learning: Active construction of physically-plausible perturbations
Ajay Mandlekar, Yuke Zhu, Animesh Garg, Li Fei-Fei, and Silvio Savarese · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Cited alongside, same era.
CAD2RL: real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2018
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Later among the works it cites.
Structured control nets for deep reinforcement learning
Mario Srouji, Jian Zhang, and Ruslan Salakhutdinov · 2018
Later among the works it cites.
Vizdoom competitions: playing doom from pixels
Marek Wydmuch, Michał Kempka, and Wojciech Jaśkowski · 2018
Later among the works it cites.
A dissection of overfitting and generalization in continuous reinforcement learning
Amy Zhang, Nicolas Ballas, and Joelle Pineau · 2018
Later among the works it cites.
Natural environment benchmarks for reinforcement learning
Amy Zhang, Yuxin Wu, and Joelle Pineau · 2018
Later among the works it cites.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio · 2018
Later among the works it cites.
Deep reinforcement learning on a budget: 3d control and reasoning without a supercomputer
Edward Beeching, Christian Wolf, Jilles Dibangoye, and Olivier Simonin · 2019
Closest in time.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Closest in time.
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, DJ Strouse, Zafarali Ahmed, Hugo Larochelle, Matthew Botvinick, Sergey Levine, and Yoshua Bengio · 2019
Closest in time.
Obstacle tower: A generalization challenge in vision, control, and planning
Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper, Hunter Henry, Adam Crespi, Julian Togelius, and Danny Lange · 2019
Closest in time.
Rogue-gym: A new challenge for generalization in reinforcement learning
Yuji Kanagawa and Tomoyuki Kaneko · 2019
Closest in time.
Towards understanding regularization in batch normalization
Ping Luo, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.
Deep reinforcement learning with relational inductive biases
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, Murray Shanahan, Victoria Langston, Razvan Pascanu, Matthew Botvinick, Oriol Vinyals, and Peter Battaglia · 2019
Closest in time.
Investigating generalisation in continuous deep reinforcement learning
Chenyang Zhao, Olivier Siguad, Freek Stulp, and Timothy M Hospedales · 2019
Closest in time.