Fetching the paper…
Reading the bibliography…
Many real world stochastic control problems suffer from the "curse of dimensionality".
Dynamic Programming
Richard Ernest Bellman · 1957
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Optimal control of execution costs
Dimitris Bertsimas and Andrew W. Lo · 1998
Earlier work this paper cites.
Optimal control of execution costs for portfolios
Dimitris Bertsimas, Andrew W. Lo, and P. Hummel · 1999
Earlier work this paper cites.
Learning deep architectures for AI
Y. Bengio · 2009
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the Curses of Dimensionality
Warren B. Powell · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Cited alongside, same era.
Benchmarking a scalable approximation dynamic programming algorithm for stochastic control of multidimensional energy storage problems
Daniel F. Salas and Warren B. Powell · 2013
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, et al · 2015
Adam: a method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Later among the works it cites.
An approximate dynamic programming algorithm for monotone value functions
Daniel R. Jiang and Warren B. Powell · 2015
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, et al · 2016
Closest in time.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, and David Silver Daan Wierstra Tom Erez, Yuval Tassa · 2016
Closest in time.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Closest in time.
Benchmarking deep reinforcement learning for continuous control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Closest in time.