Fetching the paper…
Reading the bibliography…
In this paper we show how risk-averse reinforcement learning can be used to hedge options.
The pricing of options and corporate liabilities
Fischer Black and Myron Scholes. 1973 · 1973
Earlier work this paper cites.
The variance of discounted Markov decision processes
Matthew J. Sobel. 1982 · 1982
Earlier work this paper cites.
Introduction to Reinforcement Learning (1st ed.)
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Gradient descent for general reinforcement learning. In NeurIPS . 968–974
Leemon C Baird III and Andrew W Moore. 1999 · 1999
Earlier work this paper cites.
Learning to trade via direct reinforcement
John Moody and Matthew Saffell. 2001 · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning. In ICML . 267–274
Sham Kakade and John Langford. 2002 · 2002
Earlier work this paper cites.
Nonparametric Return Distribution Approximation for Reinforcement Learning. In ICML
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka. 2010 · 2010
Earlier work this paper cites.
Policy Gradients with Variance Related Risk Criteria
Dotan Di Castro, Aviv Tamar, and Shie Mannor. 2012 · 2012
Earlier work this paper cites.
Criticism of the Black-Scholes Model: But Why is It Still Used?
Orhun Hakan Yalincak. 2012 · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Adaptive Step-Size for Policy Gradient Methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta. 2013 · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley. 2013 · 2013
Cited alongside, same era.
Variance adjusted actor critic algorithms
Aviv Tamar and Shie Mannor. 2013 · 2013
Cited alongside, same era.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Risk-averse reinforcement learning for algorithmic trading. In CIFEr . 391–398
Y. Shen, R. Huang, C. Yan, and K. Obermayer. 2014 · 2014
Cited alongside, same era.
Trust Region Policy Optimization. In ICML , Vol. 37. 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Policy Gradient for Coherent Risk Measures
Adaptive batch size for safe policy gradients. In NeurIPS . 3591–3600
Matteo Papini, Matteo Pirotta, and Marcello Restelli. 2017 · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Later among the works it cites.
Curve Dynamics with Artificial Neural Networks
Alexei Kondratyev. 2018 · 2018
Later among the works it cites.
OpenAI Five
OpenAI. 2018 · 2018
Later among the works it cites.
Risk-Averse Trust Region Optimization for Reward-Volatility Reduction
Lorenzo Bisi, Luca Sabbioni, Edoardo Vittori, Matteo Papini, and Marcello Restelli. 2019 · 2019
Later among the works it cites.
Deep hedging
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor. 2015 · 2015
Cited alongside, same era.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
L Prashanth and Mohammad Ghavamzadeh. 2016 · 2016
Cited alongside, same era.
Learning the variance of the reward-to-go
Aviv Tamar, Dotan Di Castro, and Shie Mannor. 2016 · 2016
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone. 2017 · 2017
Cited alongside, same era.
Emergence of Locomotion Behaviours in Rich Environments
Nicolas Heess, Dhruva TB, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, S. M. Ali Eslami, Martin A. Riedmiller, and David Silver. 2017 · 2017
Cited alongside, same era.
Optimal delta hedging for options
John Hull and Alan White. 2017 · 2017
Cited alongside, same era.
Hans Buehler, Lukas Gonon, Josef Teichmann, and Ben Wood. 2019 · 2019
Later among the works it cites.
Deep Hedging of Derivatives Using Reinforcement Learning
Jay Cao, Jacky Chen, John C. Hull, and Zissis Poulos. 2019 · 2019
Later among the works it cites.
The QLBS Q-Learner goes NuQLear: fitted Q iteration, inverse RL, and option portfolios
Igor Halperin. 2019 · 2019
Later among the works it cites.
Dynamic replication and hedging: A reinforcement learning approach
Petter N Kolm and Gordon Ritter. 2019 · 2019
Later among the works it cites.
The Market Generator
Alexei Kondratyev and Schwarz Christian. 2019 · 2019
Later among the works it cites.
Qlbs: Q-learner in the black-scholes (-merton) worlds
Igor Halperin. 2020 · 2020
Closest in time.