Fetching the paper…
Reading the bibliography…
Three major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment.
A general class of coefficients of divergence of one distribution from another
Syed Mumtaz Ali and Samuel D Silvey · 1966
Earlier work this paper cites.
Information-type measures of difference of probability distributions and indirect observation
Imre Csiszár · 1967
Earlier work this paper cites.
Multinomial goodness-of-fit tests
Noel Cressie and Timothy RC Read · 1984
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Gaussian processes for machine learning, vol. 1, 2006
Carl Edward Rasmussen and CK Williams · 2006
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Distributionally robust Markov decision processes
Huan Xu and Shie Mannor · 2010
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Robust Markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Reinforcement learning in robust Markov decision processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L Jeff Hong · 2013
Earlier work this paper cites.
Model predictive control
Eduardo F Camacho and Carlos Bordons Alba · 2013
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Aharon Ben-Tal, Dick Den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Earlier work this paper cites.
Distributionally robust counterpart in markov decision processes
Pengqian Yu and Huan Xu · 2015
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
Paul Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Robust mdps with k-rectangular uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2016
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Mutual alignment transfer learning
Markus Wulfmeier, Ingmar Posner, and Pieter Abbeel · 2017
Cited alongside, same era.
Reinforcement learning under model mismatch
Aurko Roy, Huan Xu, and Sebastian Pokutta · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Cited alongside, same era.
Distributionally robust stochastic programming
Alexander Shapiro · 2017
Cited alongside, same era.
Model-based reinforcement learning via meta-policy optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Efficient planning in large MDPs with weak linear function approximation
Roshan Shariff and Csaba Szepesvári · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2020
Later among the works it cites.
A bayesian approach to robust reinforcement learning
Esther Derman, Daniel Mankowitz, Timothy Mann, and Shie Mannor · 2020
Later among the works it cites.
Robust deep reinforcement learning against adversarial perturbations on state observations
Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan Liu, Duane Boning, and Cho-Jui Hsieh · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample-efficient reinforcement learning via difference models
Divyam Rastogi, Ivan Koryakovskiy, and Jens Kober · 2018
Cited alongside, same era.
Adversarially robust optimization with Gaussian processes
Ilija Bogunovic, Jonathan Scarlett, Stefanie Jegelka, and Volkan Cevher · 2018
Cited alongside, same era.
Gaussian processes and kernel methods: A review on connections and equivalences
Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, and Bharath K Sriperumbudur · 2018
Cited alongside, same era.
Soft-robust actor-critic policy-gradient
Esther Derman, Daniel J Mankowitz, Timothy A Mann, and Shie Mannor · 2018
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
John Duchi and Hongseok Namkoong · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Sebastian Curi, Felix Berkenkamp, and Andreas Krause · 2020
Later among the works it cites.
Distributionally robust Bayesian optimization
Johannes Kirschner, Ilija Bogunovic, Stefanie Jegelka, and Andreas Krause · 2020
Later among the works it cites.
Distributionally robust Bayesian quadrature optimization
Thanh Nguyen, Sunil Gupta, Huong Ha, Santu Rana, and Svetha Venkatesh · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
Sham Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi, and Wen Sun · 2020
Later among the works it cites.
Sample-efficient cross-entropy method for real-time planning
Cristina Pinneri, Shambhuraj Sawant, Sebastian Blaes, Jan Achterhold, Joerg Stueckler, Michal Rolinek, and Georg Martius · 2020
Later among the works it cites.
An experimental design perspective on model-based reinforcement learning
Viraj Mehta, Biswajit Paria, Jeff Schneider, Stefano Ermon, and Willie Neiswanger · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Later among the works it cites.
Wenhao Yang, Liangyu Zhang, and Zhihua Zhang · 2021
Later among the works it cites.
Robust reinforcement learning using least squares policy iteration with provable performance guarantees
Kishan Panaganti Badrinath and Dileep Kalathil · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Yue Wang and Shaofeng Zou · 2021
Later among the works it cites.
Combining pessimism with optimism for robust and efficient model-based deep reinforcement learning
Sebastian Curi, Ilija Bogunovic, and Andreas Krause · 2021
Later among the works it cites.
Optimal order simple regret for gaussian process bandits
Sattar Vakili, Nacime Bouziani, Sepehr Jalali, Alberto Bernacchia, and Da-shan Shiu · 2021
Later among the works it cites.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Later among the works it cites.
Robust reinforcement learning using offline data
Kishan Panaganti, Zaiyan Xu, Dileep Kalathil, and Mohammad Ghavamzadeh · 2022
Later among the works it cites.
When should we prefer offline reinforcement learning over behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2022
Later among the works it cites.
Near-optimal policy identification in active reinforcement learning
Xiang Li, Viraj Mehta, Johannes Kirschner, Ian Char, Willie Neiswanger, Jeff Schneider, Andreas Krause, and Ilija Bogunovic · 2023
Closest in time.