Fetching the paper…
Reading the bibliography…
In the classical Reinforcement Learning (RL) setting, one aims to find a policy that maximizes its expected return.
Zur theorie der gesellschaftsspiele
J. v. Neumann · 1928
Earlier work this paper cites.
On general minimax theorems
M. Sion · 1958
Earlier work this paper cites.
Stochastic and shortest path games: theory and algorithms
S. D. Patek · 1997
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Coherent measures of risk
P. Artzner, F. Delbaen, J.-M. Eber, and D. Heath · 1999
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Market structure and equilibrium
H. Von Stackelberg · 2010
Earlier work this paper cites.
Variance adjusted actor critic algorithms
A. Tamar and S. Mannor · 2013
Earlier work this paper cites.
Lectures on stochastic programming: modeling and theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Y. Chow, A. Tamar, S. Mannor, and M. Pavone · 2015
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum markov games
J. Perolat, B. Scherrer, B. Piot, and O. Pietquin · 2015
Earlier work this paper cites.
Policy gradient in lipschitz markov decision processes
M. Pirotta, M. Restelli, and L. Bascetta · 2015
Earlier work this paper cites.
Policy gradient for coherent risk measures
A. Tamar, Y. Chow, M. Ghavamzadeh, and S. Mannor · 2015
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Earlier work this paper cites.
BEGAN: boundary equilibrium generative adversarial networks
D. Berthelot, T. Schumm, and L. Metz · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Adversarial machine learning at scale
A. Kurakin, I. J. Goodfellow, and S. Bengio · 2017
Cited alongside, same era.
Adversarially robust policy learning: Active construction of physically-plausible perturbations
A. Mandlekar, Y. Zhu, A. Garg, L. Fei-Fei, and S. Savarese · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Cited alongside, same era.
Action robust reinforcement learning and applications in continuous control
C. Tessler, Y. Efroni, and S. Mannor · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Implicit learning dynamics in stackelberg games: Equilibria characterization, convergence analysis, and empirical study
T. Fiez, B. Chasnov, and L. Ratliff · 2020
Later among the works it cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Later among the works it cites.
Robust reinforcement learning via adversarial training with langevin dynamics
P. Kamalaruban, Y.-T. Huang, Y.-P. Hsieh, P. Rolland, C. Shi, and V. Cevher · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning
D. Baek, M. Hwang, H. Kim, and D.-S. Kwon · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
W. Dabney, G. Ostrovski, D. Silver, and R. Munos · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Cited alongside, same era.
Optimal approximation of piecewise smooth functions using deep relu neural networks
P. Petersen and F. Voigtländer · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Cited alongside, same era.
Being optimistic to be conservative: Quickly learning a cvar policy
R. Keramati, C. Dann, A. Tamkin, and E. Brunskill · 2020
Later among the works it cites.
Model-based adversarial meta-reinforcement learning
Z. Lin, G. Thomas, G. Yang, and T. Ma · 2020
Later among the works it cites.
Learning domain randomization distributions for training robust locomotion policies
M. Mozian, J. C. G. Higuera, D. Meger, and G. Dudek · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
A. Rajeswaran, I. Mordatch, and V. Kumar · 2020
Later among the works it cites.
Improving robustness via risk averse distributional reinforcement learning
R. Singh, Q. Zhang, and Y. Chen · 2020
Later among the works it cites.
Two steps to risk sensitivity
C. Gagne and P. Dayan · 2021
Closest in time.
Automatic risk adaptation in distributional reinforcement learning
F. Schubert, T. Eimer, B. Rosenhahn, and M. Lindauer · 2021
Closest in time.
Tianshou: A highly modularized deep reinforcement learning library
J. Weng, H. Chen, D. Yan, K. You, A. Duburcq, M. Zhang, H. Su, and J. Zhu · 2021
Closest in time.
Safe distributional reinforcement learning
J. Zhang and P. Weng · 2021
Closest in time.