Fetching the paper…
Reading the bibliography…
In this thesis, we develop a comprehensive account of the expressive power, modelling efficiency, and performance advantages of so-called trading agents (i.e., Deep Soft Recurrent Q-Network (DSRQN) and Mixture of Score Machines (MSM)), based on both traditional system identification (model-based approach) as well as on context-independent agents (model-free approach).
“Augmented MVDR spectrum-based frequency estimation for unbalanced power systems”
Yili Xia and Danilo Mandic · 1926
Earlier work this paper cites.
“Asynchronous methods for deep reinforcement learning”
Volodymyr Mnih et al · 1937
Earlier work this paper cites.
“Portfolio selection”
Harry Markowitz · 1952
Earlier work this paper cites.
“Dynamic programming”
Richard Bellman · 1957
Earlier work this paper cites.
“Efficient capital markets: A review of theory and empirical work”
Eugene Fama · 1970
Earlier work this paper cites.
“Portfolio theory and capital markets”
William Sharpe and WF Sharpe · 1970
Earlier work this paper cites.
“Nonlinear regulator theory and an inverse optimal control problem”
P Moylan and B Anderson · 1973
Earlier work this paper cites.
“An adaptive optimal controller for discrete-time Markov environments”
Ian Witten · 1977
Earlier work this paper cites.
“Practical optimization”
Philip Gill, Walter Murray and Margaret Wright · 1981
Earlier work this paper cites.
“Neural networks and physical systems with emergent collective computational abilities”
John Hopfield · 1982
Earlier work this paper cites.
“Mean-variance versus direct utility maximization”
Yoram Kroll, Haim Levy and Harry Markowitz · 1984
Earlier work this paper cites.
“Generalized linear models”
Peter McCullagh · 1984
Earlier work this paper cites.
“Generalized autoregressive conditional heteroskedasticity”
Tim Bollerslev · 1986
Earlier work this paper cites.
“Approximation by superpositions of a sigmoidal function”
George Cybenko · 1989
Earlier work this paper cites.
“Learning from delayed rewards”, 1989
Christopher John Cornish Watkins · 1989
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it”
Paul Werbos · 1990
Earlier work this paper cites.
“Adaptive mixtures of local experts”
Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton · 1991
Earlier work this paper cites.
“Financial indicators and growth in a cross section of countries”
Robert King and Ross Levine · 1992
Earlier work this paper cites.
“Generating surrogate data for time series with several simultaneously measured variables”
Dean Prichard and James Theiler · 1994
Earlier work this paper cites.
“Continuous-time nonlinear signal processing: a neural network based approach for gray box identification”
R Rico-Martinez, JS Anderson and IG Kevrekidis · 1994
Earlier work this paper cites.
“Dynamic programming and optimal control”
Dimitri Bertsekas, Dimitri Bertsekas, Dimitri Bertsekas and Dimitri Bertsekas · 1995
Earlier work this paper cites.
“Convolutional networks for images, speech, and time series”
Yann LeCun and Yoshua Bengio · 1995
Earlier work this paper cites.
“Temporal difference learning and TD-Gammon”
Gerald Tesauro · 1995
Earlier work this paper cites.
“Optimal asset allocation using adaptive dynamic programming”
Ralph Neuneier · 1996
Earlier work this paper cites.
“A model of multiplicative neural responses in parietal cortex”
Emilio Salinas and LF Abbott · 1996
Earlier work this paper cites.
“A comparison of direct and model-based reinforcement learning”
Christopher Atkeson and Juan Santamaria · 1997
Earlier work this paper cites.
“Long short-term memory”
Sepp Hochreiter and J“”urgen Schmidhuber · 1997
Earlier work this paper cites.
“Investment science”
David Luenberger · 1997
Earlier work this paper cites.
“A comparison between recurrent neural network architectures for digital equalization”
Jorge Ortiz-Fuentes and Mikel Forcada · 1997
Earlier work this paper cites.
“Using randomization to break the curse of dimensionality”
John Rust · 1997
Earlier work this paper cites.
“Markovian representation of stochastic processes and its application to the analysis of autoregressive moving average processes”
Hirotugu Akaike · 1998
Earlier work this paper cites.
“The generalization of the wiener-khinchin theorem”
Leon Cohen · 1998
Earlier work this paper cites.
“The vanishing gradient problem during learning recurrent neural nets and problem solutions”
Sepp Hochreiter · 1998
Earlier work this paper cites.
“Reinforcement learning for trading systems and portfolios: Immediate vs future rewards”
John Moody, Matthew Saffell, Yuansong Liao and Lizhong Wu · 1998
Earlier work this paper cites.
“Combinatorial optimization: Algorithms and complexity”
Christos Papadimitriou and Kenneth Steiglitz · 1998
Earlier work this paper cites.
“Introduction to reinforcement learning”
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
“Learning to forget: Continual prediction with LSTM”
Felix Gers, J“”urgen Schmidhuber and Fred Cummins · 1999
Earlier work this paper cites.
“Algorithms for inverse reinforcement learning.”
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
“Policy gradient methods for reinforcement learning with function approximation”
Richard Sutton, David McAllester, Satinder Singh and Yishay Mansour · 2000
Earlier work this paper cites.
“Policy gradient methods for reinforcement learning with function approximation”
Richard Sutton, David McAllester, Satinder Singh and Yishay Mansour · 2000
Earlier work this paper cites.
“An introduction to hidden Markov models and Bayesian networks”
Zoubin Ghahramani · 2001
Earlier work this paper cites.
“A builder’s guide to agent-based financial markets”
Blake LeBaron · 2001
Earlier work this paper cites.
“Recurrent neural networks for prediction: Learning algorithms, architectures and stability”
Danilo Mandic and Jonathon Chambers · 2001
Earlier work this paper cites.
“Financial volatility trading using recurrent neural networks”
Peter Tino, Christian Schittenkopf and Georg Dorffner · 2001
Cited alongside, same era.
“Deep blue”
Murray Campbell, A Hoane and Feng-hsiung Hsu · 2002
Cited alongside, same era.
“Effective reinforcement learning for mobile robots”
William Smart and L Kaelbling · 2002
Cited alongside, same era.
“A neural probabilistic language model”
Yoshua Bengio, R“’ejean Ducharme, Pascal Vincent and Christian Jauvin · 2003
Cited alongside, same era.
“Econometric analysis”
William Greene · 2003
Cited alongside, same era.
“Convex optimization”
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
“Policy gradient reinforcement learning for fast quadrupedal locomotion”
“Continuous control with deep reinforcement learning”
Timothy Lillicrap et al · 2015
Later among the works it cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Later among the works it cites.
“Integrating learning and planning”, 2015
David Silver · 2015
Later among the works it cites.
“Introduction to reinforcement learning”, 2015
David Silver · 2015
Later among the works it cites.
“Markov decision processes”, 2015
David Silver · 2015
Later among the works it cites.
“Model-free control”, 2015
David Silver · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nate Kohl and Peter Stone · 2004
Cited alongside, same era.
“A generalized normalized gradient descent algorithm”
Danilo Mandic · 2004
Cited alongside, same era.
“Gaussian processes in machine learning”
Carl Rasmussen · 2004
Cited alongside, same era.
“Analysis of financial time series”
Ruey Tsay · 2005
Cited alongside, same era.
“Pairs trading: Performance of a relative-value arbitrage rule”
Evan Gatev, William Goetzmann and K Rouwenhorst · 2006
Cited alongside, same era.
“Pattern recognition and machine learning”
Nasser Nasrabadi · 2007
Cited alongside, same era.
David Silver · 2015
Later among the works it cites.
“A signal processing perspective on financial engineering”
Yiyong Feng and Daniel Palomar · 2016
Later among the works it cites.
“Uncertainty in deep learning”
Yarin Gal · 2016
Later among the works it cites.
“A theoretically grounded application of dropout in recurrent neural networks”
Yarin Gal and Zoubin Ghahramani · 2016
Later among the works it cites.
“Improving PILCO with Bayesian neural network dynamics models”
Yarin Gal, Rowan McAllister and Carl Rasmussen · 2016
Later among the works it cites.
“Deep learning”
Ian Goodfellow, Yoshua Bengio, Aaron Courville and Yoshua Bengio · 2016
Later among the works it cites.
JB Heaton, NG Polson and Jan Witte · 2016
Later among the works it cites.
“Stochastic financial models”
Douglas Kennedy · 2016
Later among the works it cites.
“End-to-end training of deep visuomotor policies”
Sergey Levine, Chelsea Finn, Trevor Darrell and Pieter Abbeel · 2016
Later among the works it cites.
“State of the art control of atari games using shallow reinforcement learning”
Yitao Liang, Marlos Machado, Erik Talvitie and Michael Bowling · 2016
Later among the works it cites.
“Policy gradient algorithms for asset allocation problem”, 2016
Pierpaolo Necchi · 2016
Later among the works it cites.
“Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach”
Shamim Nemati, Mohammad Ghassemi and Gari Clifford · 2016
Later among the works it cites.
“AlphaGo: Mastering the ancient game of Go with Machine Learning”
David Silver and Demis Hassabis · 2016
Later among the works it cites.
“A deep learning framework for financial time series using stacked autoencoders and long-short term memory”
Wei Bao, Jun Yue and Yulei Rao · 2017
Later among the works it cites.
“Deep direct reinforcement learning for financial signal representation and trading”
Yue Deng et al · 2017
Later among the works it cites.
“Deep learning for finance: Deep portfolios”
JB Heaton, NG Polson and Jan Witte · 2017
Later among the works it cites.
“A deep reinforcement learning framework for the financial portfolio management problem”
Zhengyao Jiang, Dixing Xu and Jinjun Liang · 2017
Later among the works it cites.
“Financial time series prediction using deep learning”
Ariel Navon and Yosi Keller · 2017
Later among the works it cites.
“JPMorgan develops robot to execute trades”, 2017
Laura Noonan · 2017
Later among the works it cites.
“Commission models”, 2017
Quantopian · 2017
Later among the works it cites.
“An essay on financial information in the era of computerization”
Christophe Schinckus · 2017
Later among the works it cites.
“Signal processing for finance, economics, and marketing: Concepts, framework, and big data applications”
Xiao-Ping Zhang and Fang Wang · 2017
Later among the works it cites.
“Fourier Policy Gradients”
Matthew Fellows, Kamil Ciosek and Shimon Whiteson · 2018
Later among the works it cites.
“Euro STOXX 50 index”, 2018
Investopedia · 2018
Later among the works it cites.
“Law of supply and demand”, 2018
Investopedia · 2018
Later among the works it cites.
“Liquidity”, 2018
Investopedia · 2018
Later among the works it cites.
“Market value”, 2018
Investopedia · 2018
Later among the works it cites.
“Slippage”, 2018
Investopedia · 2018
Later among the works it cites.
“Standard & Poor’s 500 index - S&P 500”, 2018
Investopedia · 2018
Later among the works it cites.
“Transaction costs”, 2018
Investopedia · 2018
Later among the works it cites.
“Advanced signal processing: Linear stochastic processes”, 2018
Danilo. Mandic · 2018
Later among the works it cites.
“Introduction to estimation theory”, 2018
Danilo. Mandic · 2018
Later among the works it cites.
“Market Making via Reinforcement Learning”
Thomas Spooner, John Fearnley, Rahul Savani and Andreas Koukorinis · 2018
Later among the works it cites.
“London Stock Exchange”, 2018
Wikipedia · 2018
Later among the works it cites.
“NASDAQ”, 2018
Wikipedia · 2018
Later among the works it cites.
“New York Stock Exchange”, 2018
Wikipedia · 2018
Later among the works it cites.