Fetching the paper…
Reading the bibliography…
This paper presents a discrete-time option pricing model that is rooted in Reinforcement Learning (RL), and more specifically in the famous Q-Learning method of RL.
H. Robbins and S. Monro, “A Stochastic Approximation Method”, Ann. Math. Statistics, 22, 400-407, 1951
1951
Earlier work this paper cites.
H. Markowitz, Portfolio Selection: efficient diversification of investment
1959
Earlier work this paper cites.
L. Bachelier, “Théorie de la spéculation”, Annales Scientifiques de I’École Normale Supérieure
1964
Earlier work this paper cites.
P. Samuelson, “Rational theory of warrant pricing”, Industrial Management Review
1965
Earlier work this paper cites.
F. Black and M. Scholes, “The Pricing of Options and Corporate Liabilities”, Journal of Political Economy, Vol. 81(3), 637-654, 1973
1973
Earlier work this paper cites.
R. Merton, “Theory of Rational Option Pricing”, Bell Journal of Economics and Management Science, Vol.4(1), 141-183, 1974
1974
Earlier work this paper cites.
H. Föllmer and M. Schweizer, “Hedging by Sequential Regression: an Introduction to the Mathematics of Option Trading”, ASTIN Bulletin
1989
Earlier work this paper cites.
C.J. Watkins, Learning from Delayed Rewards
1989
Earlier work this paper cites.
C.J. Watkins and P. Dayan, “Q-Learning”, Machine Learning, 8(3-4), 179-192, 1992
1992
Earlier work this paper cites.
M. Schweizer, “Variance-Optimal Hedging in Discrete Time”, Mathematics of Operations Research, 20, 1-32, 1995
1995
Cited alongside, same era.
P. Wilmott, Derivatives: The Theory and Practive of Financial Engineering
1998
Cited alongside, same era.
J.C. Duan and J.G. Simonato, “American Option Pricing Under GARCH by a Markov Chain Approximation”, Journal of Economic Dynamics and Control, Vol. 25, (2001), pp. 1689-1718
2001
Cited alongside, same era.
F.A. Longstaff and E.S. Schwartz, “Valuing American Options by Simulation - a Simple Least-Square Approach”, The Review of Financial Studies, Vol. 14(1), 113-147, 2001
2001
Cited alongside, same era.
M. Potters, J.P. Bouchaud, and D. Sestovic, “Hedged Monte Carlo: Low Variance Derivative Pricing with Objective Probabilities”, Physica A, vol. 289, 517-525, 2001
2001
Cited alongside, same era.
A.J. Grau, “Applications of Least-Square Regressions to Pricing and Hedging of Financial Derivatives”, PhD. thesis, Technische Universit”at München, 2007
2007
Later among the works it cites.
A. Gosavi, “Finite Horizon Markov Control with One-Step Variance Penalties”, Conference Proceedings of the Allerton Conferences, Allerton, IL, 2010
2010
Later among the works it cites.
H. van Hasselt, ”Double Q-Learning”, Advances in Neural Information Processing Systems, 2010 (http://papers.nips.cc/paper/3964-double-q-learning.pdf)
2010
Later among the works it cites.
A.Petrelli et al, “Optimal Dynamic Hedging of Equity Options: Residual-Risks, Transaction-Costs, & Conditioning”, working paper
2010
Later among the works it cites.
C. Elkan, “Reinforcement Learning with a Bilinear Q Function”, working paper (2011)
2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Ernst, P. Geurts, and L. Wehenkel, “Tree-Based Batch Model Reinforcement Learning”, Journal of Machine Learning Research, 6, 405-556, 2005
2005
Cited alongside, same era.
J. Lim, “A Numerical Algorithm for Indifference Pricing in Incomplete Markets”, working paper, University of Texas, 2005
2005
Cited alongside, same era.
S.A. Murphy, “A Generalization Error for Q-Learning”, Journal of Machine Learning Research, 6, 1073-1097, 2005
2005
Cited alongside, same era.
A. Cerný and J. Kallsen, “Hedging by Sequential Regression Revisited”, Working paper, City University London and TU München, 2007
2007
Cited alongside, same era.
https://github.com/openai/gym
Cited in the paper.
“With four parameters I can fit an elephant, and with five I can make him wiggle his trunk.” (John von Neumann)
Cited in the paper.
R. Fonteneau, “Contributions to Batch Mode Reinforcement Learning”, Ph.D. thesis, University of Li’ege, 2011
2011
Later among the works it cites.
A. Gosavi, “Solving Markov Decision Processes via Simulation”, in Handbook of Simulation Optimization
2014
Later among the works it cites.
R. S. Sutton and A. G. Barto, “Reinforcement Learning: An Introduction”, second edition, MIT, 2018
2018
Closest in time.
I. Halperin, “The QLBS Q-Learner Goes NuQLear: Fitted Q Iteration, Inverse RL, and Option Portfolios”, Quantitative Finance
2019
Closest in time.