Fetching the paper…
Reading the bibliography…
Agents trained by reinforcement learning (RL) often fail to generalize beyond the environment they were trained in, even when presented with new scenarios that seem similar to the training environment.
Efficient reinforcement learning in factored mdps
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y. Ng · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael Kearns, and John Langford · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y. Ng · 2005
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Using bisimulation for policy transfer in mdps
Pablo Samuel Castro and Doina Precup · 2010
Earlier work this paper cites.
Bayesian multi-task reinforcement learning
Alessandro Lazaric and Mohammad Ghavamzadeh · 2010
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J. Kappen · 2012
Earlier work this paper cites.
Sample complexity of multi-task reinforcement learning
Emma Brunskill and Lihong Li · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V. Todorov, and Sham M Kakade · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver et al · 2017
Cited alongside, same era.
Generalization and regularization in DQN
Jesse Farebrother, Marlos C. Machado, and Michael Bowling · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Composable deep reinforcement learning for robotic manipulation
T. Haarnoja, V. Pong, A. Zhou, M. Dalal, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
PAC reinforcement learning with an imperfect model
Nan Jiang · 2018
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F. Yang · 2020
Later among the works it cites.
Instance-based generalization in reinforcement learning
Martin Bertran, Natalia Martinez, Mariano Phielipp, and Guillermo Sapiro · 2020
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman · 2020
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S. Du, Sham M. Kakade, Ruosong Wang, and Lin F. Yang · 2020
Later among the works it cites.
Agnostic q-learning with function approximation in deterministic systems: Near-optimal bounds on approximation error and sample complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Nichol, Vicki Pfau, Christopher Hesse, Oleg Klimov, and John Schulman · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
A dissection of overfitting and generalization in continuous reinforcement learning
Amy Zhang, Nicolas Ballas, and Joelle Pineau · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Ignasi Clavera, Anusha Nagabandi, Simin Liu, Ronald S. Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Cited alongside, same era.
Provably efficient q-learning with function approximation via distribution shift error checking oracle
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Cited alongside, same era.
Simon S Du, Jason D Lee, Gaurav Mahajan, and Ruosong Wang · 2020
Later among the works it cites.
Provably convergent policy gradient methods for model-agnostic meta-reinforcement learning
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in RL with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2020
Later among the works it cites.
Invariant policy optimization: Towards stronger generalization in reinforcement learning
Anoopkumar Sonar, Vincent Pacelli, and Anirudha Majumdar · 2020
Later among the works it cites.
On the global optimality of model-agnostic meta-learning
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S. Du, Lin F. Yang, and Ruslan Salakhutdinov · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G Bellemare · 2021
Closest in time.
Metrics and continuity in reinforcement learning
Charline Le Lan, Marc G. Bellemare, and Pablo Samuel Castro · 2021
Closest in time.
On query-efficient planning in mdps under linear realizability of the optimal state-value function
Gellert Weisz, Philip Amortila, Barnabas Janzer, Yasin Abbasi-Yadkori, Nan Jiang, and Csaba Szepesvari · 2021
Closest in time.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2021
Closest in time.