Fetching the paper…
Reading the bibliography…
We study zero-shot generalization in reinforcement learning-optimizing a policy on a set of training tasks to perform well on a similar but unseen test task.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 1910
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Jan Beirlant, Edward J Dudewicz, László Györfi, Edward C Van der Meulen, et al · 1997
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Harshinder Singh, Neeraj Misra, Vladimir Hnizdo, Adam Fedorowicz, and Eugene Demchuk · 2003
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 1
Dimitri Bertsekas · 2012
Earlier work this paper cites.
Texplore: real-time sample-efficient reinforcement learning for robots
Todd Hester and Peter Stone · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Learners that use little information
Raef Bassily, Shay Moran, Ido Nachum, Jonathan Shafer, and Amir Yehudayoff · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Earlier work this paper cites.
Generalization and regularization in dqn
Jesse Farebrother, Marlos C Machado, and Michael Bowling · 2018
Earlier work this paper cites.
Action schema networks: Generalised policies with deep learning
Sam Toyer, Felipe Trevizan, Sylvie Thiébaux, and Lexing Xie · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Instance-based generalization in reinforcement learning
Martin Bertran, Natalia Martinez, Mariano Phielipp, and Guillermo Sapiro · 2020
Cited alongside, same era.
Phasic policy gradient
Karl W Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2021
Later among the works it cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan P Adams, and Sergey Levine · 2021
Later among the works it cites.
Prioritized level replay
Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
Domain adversarial reinforcement learning
Bonnie Li, Vincent François-Lavet, Thang Doan, and Joelle Pineau · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Craig Boutilier, Chih-wei Hsu, Branislav Kveton, Martin Mladenov, Csaba Szepesvari, and Manzil Zaheer · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman · 2020
Cited alongside, same era.
Comparison of hand follower and dead-end filler algorithm in solving perfect mazes
YF Hendrawan · 2020
Cited alongside, same era.
The impact of non-stationarity on generalisation in deep reinforcement learning
Maximilian Igl, Gregory Farquhar, Jelena Luketina, Wendelin Boehmer, and Shimon Whiteson · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Cited alongside, same era.
Deep reinforcement and infomax learning
Bogdan Mazoure, Remi Tachet des Combes, Thang Long Doan, Philip Bachman, and R Devon Hjelm · 2020
Cited alongside, same era.
An intrinsically-motivated approach for learning highly exploring and fast mixing policies
Mirco Mutti and Marcello Restelli · 2020
Cited alongside, same era.
Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate
Mirco Mutti, Lorenzo Pratissoli, and Marcello Restelli · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Roberta Raileanu and Rob Fergus · 2021
Later among the works it cites.
Automatic data augmentation for generalization in reinforcement learning
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Later among the works it cites.
Invariant policy optimization: Towards stronger generalization in reinforcement learning
Anoopkumar Sonar, Vincent Pacelli, and Anirudha Majumdar · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Neuro-algorithmic policies enable fast combinatorial generalization
Marin Vlastelica, Michal Rolínek, and Georg Martius · 2021
Later among the works it cites.
Meta-learning with fewer tasks through task interpolation
Huaxiu Yao, Linjun Zhang, and Chelsea Finn · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
The importance of non-markovianity in maximum state entropy exploration
Mirco Mutti, Riccardo De Santi, and Marcello Restelli · 2022
Later among the works it cites.
Regularization guarantees generalization in bayesian reinforcement learning through algorithmic stability
Aviv Tamar, Daniel Soudry, and Ev Zisselman · 2022
Later among the works it cites.
Outracing champion gran turismo drivers with deep reinforcement learning
Peter R Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, et al · 2022
Later among the works it cites.