Fetching the paper…
Reading the bibliography…
This paper gives specific divergence examples of value-iteration for several major Reinforcement Learning and Adaptive Dynamic Programming algorithms, when using a function approximator for the value function.
R. E. Bellman, Dynamic Programming . Princeton, NJ, USA: Princeton University Press, 1957
1957
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine Learning , vol. 3, pp. 9–44, 1988
1988
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, Cambridge University, 1989
1989
Earlier work this paper cites.
P. J. Werbos, “Approximating dynamic programming for real-time control and neural modeling.” Handbook of Intelligent Control, editors White and Sofge, Chapter 13 , pp. 493–525, 1992
1992
Earlier work this paper cites.
G. Rummery and M. Niranjan, “On-line q-learning using connectionist systems,” Tech. Rep. Technical Report CUED/F-INFENG/TR 166, Cambridge University Engineering Department , 1994
1994
Earlier work this paper cites.
L. C. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in International Conference on Machine Learning , 1995, pp. 30–37
1995
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “An analysis of temporal-difference learning with function approximation,” IEEE Transactions on Automatic Control, Tech. Rep., 1996
1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “Feature-based methods for large scale dynamic programming,” Machine Learning , vol. 22, no. 1-3, pp. 59–94, 1996
1996
Cited alongside, same era.
D. Prokhorov and D. Wunsch, “Adaptive critic designs,” IEEE Transactions on Neural Networks , vol. September, pp. 997–1007, 1997
1997
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . Cambridge, Massachussetts, USA: The MIT Press, 1998
1998
Cited alongside, same era.
P. J. Werbos, “Stable adaptive control using new critic designs,” eprint arXiv:adap-org/9810001 , 1998
1998
Cited alongside, same era.
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems 12 , vol. 12, 2000, pp. 1057–1063
2000
F.-Y. Wang, H. Zhang, and D. Liu, “Adaptive dynamic programming: An introduction,” IEEE Computational Intelligence Magazine , pp. 39–47, 2009
2009
Later among the works it cites.
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th Annual International Conference on Machine Learning , ser. ICML ’09. New York, NY, USA: ACM, 2009, pp. 993–1000
2009
Later among the works it cites.
H. Maei, C. Szepesvari, S. Bhatnager, D. Precup, D. Silver, and R. Sutton, “Convergent temporal-difference learning with arbitrary smooth function approximation,” in Advances in Neural Information Processing Systems (NIPS’09) . MIT Press, 2009
2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. Ferrari and R. F. Stengel, “Model-based adaptive critic designs,” Handbook of learning and approximate dynamic programming, editors Jennie Si et al. , pp. 65–96, 2004
2004
Cited alongside, same era.
2008
Cited alongside, same era.
2011
Closest in time.
2011
Closest in time.
——, “Value-gradient learning,” in Proceedings of the IEEE International Joint Conference on Neural Networks 2012 (IJCNN’12) . IEEE Press, June 2012, pp. 3062–3069
2012
Closest in time.
M. Fairbank and E. Alonso, “A comparison of learning speed and ability to cope without exploration between DHP and TD(0),” in Proceedings of the IEEE International Joint Conference on Neural Networks 2012 (IJCNN’12) . IEEE Press, June 2012, pp. 1478–1485
2012
Closest in time.