Fetching the paper…
Reading the bibliography…
Distributional Reinforcement Learning (RL) estimates return distribution mainly by learning quantile values via minimizing the quantile Huber loss function, entailing a threshold parameter often selected heuristically or via hyperparameter search, which may not generalize well and can be suboptimal.
“Robust estimation of a location parameter,”
P. J. Huber, · 1992
Earlier work this paper cites.
“Managing smile risk,”
P. S. Hagan, D. Kumar, A. Lesniewski, and D. E. Woodward, · 2002
Earlier work this paper cites.
Optimal transport: old and new
C. Villani et al., · 2009
Earlier work this paper cites.
“Robust forward algorithms via pac-bayes and laplace distributions,”
A. Noy and K. Crammer, · 2014
Earlier work this paper cites.
“A distributional perspective on reinforcement learning,”
M. G. Bellemare, W. Dabney, and R. Munos, · 2017
Earlier work this paper cites.
“Distributional reinforcement learning with quantile regression,”
W. Dabney, M. Rowland, M. Bellemare, and R. Munos, · 2018
Earlier work this paper cites.
“Implicit quantile networks for distributional reinforcement learning,”
W. Dabney, G. Ostrovski, D. Silver, and R. Munos, · 2018
Earlier work this paper cites.
“Distributed distributional deterministic policy gradients,”
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. Tb, A. Muldal, N. Heess, and T. Lillicrap, · 2018
Earlier work this paper cites.
“Fully parameterized quantile function for distributional reinforcement learning,”
D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T.-Y. Liu, · 2019
Earlier work this paper cites.
“Distributional deep reinforcement learning with a mixture of gaussians,”
Y. Choi, K. Lee, and S. Oh, · 2019
Cited alongside, same era.
“Generalized huber loss for robust learning and its efficient minimization for a robust statistics,”
K. Gokcesu and H. Gokcesu, · 2021
Cited alongside, same era.
“An alternative probabilistic interpretation of the huber loss,”
G. P. Meyer, · 2021
Cited alongside, same era.
“Gmac: A distributional perspective on actor-critic framework,”
D. W. Nam, Y. Kim, and C. Y. Park, · 2021
Cited alongside, same era.
“Robust reinforcement learning with distributional risk-averse formulation,”
P. Clavier, S. Allassonière, and E. L. Pennec, · 2022
Cited alongside, same era.
“Robust losses for learning value functions,”
A. Patterson, V. Liao, and M. White, · 2022
Later among the works it cites.
“Popo: Pessimistic offline policy optimization,”
Q. He, X. Hou, and Y. Liu, · 2022
Later among the works it cites.
“Uncertainty-aware transfer across tasks using hybrid model-based successor feature reinforcement learning,”
P. Malekzadeh, M. Hou, and K. N. Plataniotis, · 2023
Later among the works it cites.
“A unified uncertainty-aware exploration: Combining epistemic and aleatory uncertainty,”
P. Malekzadeh, M. Hou, and K. N. Plataniotis, · 2023
Later among the works it cites.
“Gamma and vega hedging using deep distributional reinforcement learning,”
J. Cao, J. Chen, S. Farghadani, J. Hull, Z. Poulos, Z. Wang, and J. Yuan, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Robust estimation and shrinkage in ultrahigh dimensional expectile regression with heavy tails and variance heterogeneity,”
J. Zhao, G. Yan, and Y. Zhang, · 2022
Cited alongside, same era.
“Point forecasting and forecast evaluation with generalized huber loss,”
R. J. Taggart, · 2022
Cited alongside, same era.
“Evaluation of point forecasts for extreme events using consistent scoring functions,”
R. Taggart, · 2022
Cited alongside, same era.
H. Tyralis, G. Papacharalampous, N. Dogulu, and K. P. Chun, · 2023
Later among the works it cites.
“Toward risk-based optimistic exploration for cooperative multi-agent reinforcement learning,”
J. Oh, J. Kim, M. Jeong, and S.-Y. Yun, · 2023
Later among the works it cites.
S. Chhachhi and F. Teng, · 2023
Later among the works it cites.