Fetching the paper…
Reading the bibliography…
The construction of replication strategies for contingent claims in the presence of risk and market friction is a key problem of financial engineering.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
The pricing of options and corporate liabilities
F. Black and M. Scholes · 1973
Earlier work this paper cites.
Theory of rational option pricing
R. C. Merton · 1973
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Option replication of contingent claims under transactions costs
S. Hodges and A. Neuberger · 1989
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
L.J Lin · 1992
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
M. L. Puterman · 1994
Earlier work this paper cites.
Variance-optimal hedging in discrete time
M. Schweizer · 1995
Earlier work this paper cites.
Analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Dynamic Asset Pricing Theory
D. Duffie · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
C. J. C. H. Watkins and P. Dayan · 2007
Earlier work this paper cites.
An empirical evaluation of thompson sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Risk-aversion in multi-armed bandits
A. Sani, A. Lazaric, and R. Munos · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Cited alongside, same era.
Scalable bayesian optimization using deep neural networks
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. A. Patwary, Prabhat, and R. Adams · 2015
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, and A. Bolton · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Later among the works it cites.
C. Riquelme, G. Tucker, and J. Snoek · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Deep hedging
H. Buehler, L. Gonon, J. Teichmann, and B. Wood · 2019
Later among the works it cites.
Deep hedging: hedging derivatives under generic market frictions using reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Risk-averse multi-armed bandit problems under mean-variance measure
V. Sattar and Z. Qing · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, and M. Lanctot · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Cited alongside, same era.
Machine learning for trading
G. Ritter · 2017
Cited alongside, same era.
H. Buehler, L. Gonon, J. Teichmann, B. Wood, B. Mohan, and J. Kochems · 2019
Later among the works it cites.
Dynamic replication and hedging: A reinforcement learning approach
P. N. Kolm and G. Ritter · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Qlbs: Q-learner in the black-scholes (-merton) worlds
I. Halperin · 2020
Closest in time.
Deep hedging of derivatives using reinforcement learning
J. Cao, J. Chen, J. Hull, and Z. Poulos · 2021
Closest in time.
Hedging of financial derivative contracts via monte carlo tree search
O. Szehr · 2021
Closest in time.