Fetching the paper…
Reading the bibliography…
In this theoretical paper we are concerned with the problem of learning a value function by a smooth general function approximator, to solve a deterministic episodic control problem in a large continuous state space.
R. Bellman, Dynamic Programming, Princeton University Press, Princeton, NJ, USA, 1957
1957
Earlier work this paper cites.
L. Sonneborn, F. V. Vleck, The bang-bang principle for linear control systems, SIAM J. Control 2 (1965) 152–159
1965
Earlier work this paper cites.
I. N. Bronshtein, K. A. Semendyayev, Handbook of Mathematics, 3rd Edition, Van Nostrand Reinhold Company, 1985, Ch. 3.2.2, pp. 372–382
1985
Earlier work this paper cites.
R. S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3 (1988) 9–44
1988
Earlier work this paper cites.
C. Watkins, Learning from delayed rewards, Ph.D. thesis, University of Cambridge, England (1989)
1989
Earlier work this paper cites.
P. J. Werbos, Backpropagation through time: What it does and how to do it, in: Proceedings of the IEEE, Vol. 78, No. 10, 1990, pp. 1550–1560
1990
Earlier work this paper cites.
C. J. C. H. Watkins, P. Dayan, Q-learning, Machine Learning 8 (1992) 279–292
1992
Earlier work this paper cites.
P. J. Werbos, Handbook of Intelligent Control, Van Nostrand, 1992
1992
Cited alongside, same era.
R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8 (1992) 229–356
1992
Cited alongside, same era.
B. A. Pearlmutter, Fast exact multiplication by the Hessian, Neural Computation 6 (1) (1994) 147–160
1994
Cited alongside, same era.
L. C. Baird, Residual algorithms: Reinforcement learning with function approximation, in: International Conference on Machine Learning, 1995, pp. 30–37
1995
Cited alongside, same era.
J. N. Tsitsiklis, B. Van Roy, An analysis of temporal-difference learning with function approximation, Tech. Rep. LIDS-P-2322 (1996)
1996
Cited alongside, same era.
K. Doya, Reinforcement learning in continuous time and space, Neural Computation 12 (1) (2000) 219–245
2000
Later among the works it cites.
R. S. Sutton, D. Mcallester, S. Singh, Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, in: Advances in Neural Information Processing Systems 12, Vol. 12, 2000, pp. 1057–1063
2000
Later among the works it cites.
C. Kwok, D. Fox, Reinforcement learning for sensing strategies, in: Proceedings of the International Confrerence on Intelligent Robots and Systems (IROS), 2004
2004
Later among the works it cites.
A. Y. Ng, H. J. Kim, M. I. Jordan, S. Sastry, Inverted autonomous helicopter flight via reinforcement learning, in: International Symposium on Experimental Robotics, MIT Press, 2004
2004
Later among the works it cites.
J. Peters, S. Schaal, Policy gradient methods for robotics, in: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2006
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, The MIT Press, Cambridge, Massachussetts, USA, 1998
1998
Cited alongside, same era.
Cited in the paper.
P. J. Werbos, Stable adaptive control using new critic designs , eprint arXiv:adap-org/9810001. URL http://xxx.lanl.gov/html/adap-org/9810001
Cited in the paper.
2006
Later among the works it cites.
R. Munos, Policy gradient in continuous time, Journal of Machine Learning Research 7 (2006) 413–427
2006
Later among the works it cites.