Fetching the paper…
Reading the bibliography…
We present a novel $Q$-learning algorithm tailored to solve distributionally robust Markov decision problems where the corresponding ambiguity set of transition probabilities for the underlying Markov decision process is a Wasserstein ball around a (possibly estimated) reference measure.
Sur les équations algébriques ayant toutes leurs racines réelles
Tiberiu Popoviciu · 1935
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On stochastic approximation
Aryeh Dvoretzky · 1956
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
Charles George Broyden · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
Roger Fletcher · 1970
Earlier work this paper cites.
A family of variable metric updates derived by variational means. mathematics of computing
Donald Goldfarb · 1970
Earlier work this paper cites.
Conditioning of quasi-Newton methods for function minimization
David F Shanno · 1970
Earlier work this paper cites.
Learning form delayed rewards
Christopher JCH Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Robust solutions to Markov decision problems with uncertain transition matrices
Laurent El Ghaoui and Arnab Nilim · 2005
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Application of Q-learning with temperature variation for bidding strategies in market based power systems
Mohammad Bagher Naghibi-Sistani, MR Akbarzadeh-Tootoonchi, MH Javidi-Dashte Bayaz, and Habib Rajabi-Mashhadi · 2006
Earlier work this paper cites.
Conditional risk mappings
Andrzej Ruszczyński and Alexander Shapiro · 2006
Earlier work this paper cites.
Model-free Q-learning designs for linear discrete-time zero-sum games with application to h-infinity control
Asma Al-Tamimi, Frank L Lewis, and Murad Abu-Khalaf · 2007
Earlier work this paper cites.
Monge-Kantorovich transportation problem and optimal couplings
Ludger Rüschendorf · 2007
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2008
Earlier work this paper cites.
Some better bounds on the variance with applications
Rajesh Sharma, Madhu Gupta, and Girish Kapoor · 2010
Earlier work this paper cites.
Markov decision processes with applications to finance
Nicole Bäuerle and Ulrich Rieder · 2011
Earlier work this paper cites.
Value-difference based exploration: adaptive control between epsilon-greedy and softmax
Michel Tokic and Günther Palm · 2011
Earlier work this paper cites.
Distributionally robust Markov decision processes
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
Robust Markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Distributionally robust convex optimization
Wolfram Wiesemann, Daniel Kuhn, and Melvyn Sim · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Distributionally robust control of constrained stochastic systems
Bart PG Van Parys, Daniel Kuhn, Paul J Goulart, and Manfred Morari · 2015
Cited alongside, same era.
Robust MDPs with k-rectangular uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2016
Cited alongside, same era.
Rectangular sets of probability measures
Alexander Shapiro · 2016
Cited alongside, same era.
A convex optimization approach to distributionally robust markov decision processes with wasserstein distance
Insoon Yang · 2017
Wasserstein distributionally robust stochastic control: A data-driven approach
Insoon Yang · 2020
Later among the works it cites.
Reinforcement learning for mean field games, with applications to economics
Andrea Angiuli, Jean-Pierre Fouque, and Mathieu Lauriere · 2021
Later among the works it cites.
Distributionally robust Markov decision processes and their connection to risk measures
Nicole Bäuerle and Alexander Glauner · 2021
Later among the works it cites.
Q-learning for distributionally robust Markov decision processes
Nicole Bäuerle and Alexander Glauner · 2021
Later among the works it cites.
Fast algorithms for l _ ∞ l\_\infty -constrained s-rectangular robust mdps
Bahram Behzadian, Marek Petrik, and Chin Pang Ho · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data-based distributionally robust stochastic optimal power flow—part i: Methodologies
Yi Guo, Kyri Baker, Emiliano Dall’Anese, Zechun Hu, and Tyler Holt Summers · 2018
Cited alongside, same era.
Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations
Peyman Mohajerin Esfahani and Daniel Kuhn · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Policy-conditioned uncertainty sets for robust markov decision processes
Andrea Tirinzoni, Marek Petrik, Xiangli Chen, and Brian Ziebart · 2018
Cited alongside, same era.
Robust optimal control using conditional risk mappings in infinite horizon
Kerem Uğurlu · 2018
Cited alongside, same era.
Data-driven risk-averse stochastic optimization with Wasserstein metric
Chaoyue Zhao and Yongpei Guan · 2018
Cited alongside, same era.
Jay Cao, Jacky Chen, John Hull, and Zissis Poulos · 2021
Later among the works it cites.
Reinforcement learning in economics and finance
Arthur Charpentier, Romuald Elie, and Carl Remlinger · 2021
Later among the works it cites.
Recent advances in reinforcement learning in finance
Ben Hambly, Renyuan Xu, and Huining Yang · 2021
Later among the works it cites.
Double deep Q-learning for optimal execution
Brian Ning, Franco Ho Ting Lin, and Sebastian Jaimungal · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Yue Wang and Shaofeng Zou · 2021
Later among the works it cites.
Wenhao Yang, Liangyu Zhang, and Zhihua Zhang · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Later among the works it cites.
Reinforcement learning algorithm for mixed mean field control games
Andrea Angiuli, Nils Detering, Jean-Pierre Fouque, and Jimin Lin · 2022
Closest in time.
Distributionally robust stochastic optimization with Wasserstein distance
Rui Gao and Anton Kleywegt · 2022
Closest in time.
Distributionally robust Q-learning
Zijian Liu, Qinxun Bai, Jose Blanchet, Perry Dong, Wei Xu, Zhengqing Zhou, and Zhengyuan Zhou · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Closest in time.
Policy gradient method for robust reinforcement learning
Yue Wang and Shaofeng Zou · 2022
Closest in time.
Robust markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clement · 2023
Closest in time.
Distributionally robust differential dynamic programming with wasserstein distance
Astghik Hakobyan and Insoon Yang · 2023
Closest in time.
Policy gradient algorithms for robust mdps with non-rectangular uncertainty sets
Mengmeng Li, Tobias Sutter, and Daniel Kuhn · 2023
Closest in time.
Bounding the difference between the values of robust and non-robust markov decision problems
Ariel Neufeld and Julian Sester · 2023
Closest in time.
Markov decision processes under model uncertainty
Ariel Neufeld, Julian Sester, and Mario Šikić · 2023
Closest in time.