Fetching the paper…
Reading the bibliography…
We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL).
Portfolio selection
Harry Markowitz · 1952
Earlier work this paper cites.
Myopia and inconsistency in dynamic utility maximization
Robert Henry Strotz · 1955
Earlier work this paper cites.
The variance of discounted Markov decision processes
Matthew J Sobel · 1982
Earlier work this paper cites.
Mean-variance hedging in continuous time
Darrell Duffie and Henry R Richardson · 1991
Earlier work this paper cites.
Common risk factors in the returns on stocks and bonds
Eugene F Fama and Kenneth R French · 1993
Earlier work this paper cites.
A nonparametric approach to pricing and hedging derivative securities via learning networks
James M Hutchinson, Andrew W Lo, and Tomaso Poggio · 1994
Earlier work this paper cites.
The econometrics of financial markets
John Y Campbell, Andrew W Lo, and A Craig MacKinlay · 1997
Earlier work this paper cites.
Investment Science
David G Luenberger · 1998
Earlier work this paper cites.
Performance functions and reinforcement learning for trading systems and portfolios
John Moody, Lizhong Wu, Yuansong Liao, and Matthew Saffell · 1998
Earlier work this paper cites.
Reinforcement learning for continuous stochastic control problems
Rémi Munos and Paul Bourgine · 1998
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Kenji Doya · 2000
Earlier work this paper cites.
Optimal dynamic portfolio selection: Multiperiod mean-variance formulation
Duan Li and Wan-Lung Ng · 2000
Earlier work this paper cites.
A study of reinforcement learning in the continuous case by the means of viscosity solutions
Rémi Munos · 2000
Earlier work this paper cites.
Variance-penalized reinforcement learning for risk-averse asset allocation
Makoto Sato and Shigenobu Kobayashi · 2000
Earlier work this paper cites.
Continuous-time mean-variance portfolio selection: A stochastic LQ framework
Xun Yu Zhou and Duan Li · 2000
Earlier work this paper cites.
Optimal execution of portfolio transactions
Robert Almgren and Neil Chriss · 2001
Cited alongside, same era.
Learning to trade via direct reinforcement
John Moody and Matthew Saffell · 2001
Cited alongside, same era.
TD algorithm for the variance of return and mean-variance reinforcement learning
Makoto Sato, Hajime Kimura, and Shibenobu Kobayashi · 2001
Cited alongside, same era.
Dynamic mean-variance portfolio selection with no-shorting constraints
Xun Li, Xun Yu Zhou, and Andrew EB Lim · 2002
Cited alongside, same era.
Mean-variance portfolio selection with random parameters in a complete market
Andrew EB Lim and Xun Yu Zhou · 2002
Cited alongside, same era.
Multiscale stochastic volatility asymptotics
Jean-Pierre Fouque, George Papanicolaou, Ronnie Sircar, and Knut Solna · 2003
Cited alongside, same era.
A reinforcement learning extension to the Almgren-Chriss framework for optimal trade execution
Dieter Hendricks and Diane Wilcox · 2014
Later among the works it cites.
Stochastic systems: Estimation, identification, and adaptive control , volume 75
Panqanamala Ramana Kumar and Pravin Varaiya · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, and Georg Ostrovski · 2015
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Cited alongside, same era.
Markowitz’s mean-variance portfolio selection with regime switching: A continuous-time model
Xun Yu Zhou and George Yin · 2003
Cited alongside, same era.
Continuous-time mean-variance portfolio selection with bankruptcy prohibition
Tomasz R Bielecki, Hanqing Jin, Stanley R Pliska, and Xun Yu Zhou · 2005
Cited alongside, same era.
Reinforcement learning for optimized trade execution
Yuriy Nevmyvaka, Yi Feng, and Michael Kearns · 2006
Cited alongside, same era.
Identification and stochastic adaptive control
Han-Fu Chen and Lei Guo · 2012
Cited alongside, same era.
Algorithmic aspects of mean–variance optimization in Markov decision processes
Shie Mannor and John N Tsitsiklis · 2013
Cited alongside, same era.
Timothy Lillicrap, Jonathan Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Later among the works it cites.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
LA Prashanth and Mohammad Ghavamzadeh · 2016
Later among the works it cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, and Marc Lanctot · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
On the policy improvement algorithm in continuous time
Saul D Jacka and Aleksandar Mijatović · 2017
Later among the works it cites.
Distributionally robust mean-variance portfolio selection with Wasserstein distances
Jose Blanchet, Lin Chen, and Xun Yu Zhou · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Exploration versus exploitation in reinforcement learning: A stochastic control approach
Haoran Wang, Thaleia Zariphopoulou, and Xun Yu Zhou · 2019
Closest in time.