Fetching the paper…
Reading the bibliography…
Reinforcement learning consists of finding policies that maximize an expected cumulative long-term reward in a Markov decision process with unknown transition probabilities and instantaneous rewards.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics
1951
Earlier work this paper cites.
Wiley for The Massachusetts Institute of Technology, 1964
R. A. Howard, Dynamic programming and Markov processes · 1964
Earlier work this paper cites.
Hermann, Paris, 1966
L. Schwartz, Théorie des distributions · 1966
Earlier work this paper cites.
S. E. Shreve and D. P. Bertsekas, “Alternative theoretical frameworks for finite horizon discrete-time stochastic optimal control,” SIAM J. on control and optimization
1978
Earlier work this paper cites.
R. Pemantle, “Nonconvergence to unstable points in urn models and stochastic approximations,” The Annals of Prob
1990
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning
1992
Earlier work this paper cites.
MIT press Cambridge, 1998
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 1998
Earlier work this paper cites.
Athena Sci., Belmont, 1999
D. P. Bertsekas, Nonlinear programming · 1999
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Adv. in neural information proc. sys
2000
Earlier work this paper cites.
Springer series in statistics New York, 2001
J. Friedman, T. Hastie, and R. Tibshirani, The elements of statistical learning · 2001
Earlier work this paper cites.
P. Vincent and Y. Bengio, “Kernel matching pursuit,” Machine Learning
2002
Cited alongside, same era.
J. Kivinen, A. J. Smola, and R. C. Williamson, “Online learning with kernels,” Trans. on Sig. Proc
2004
Cited alongside, same era.
T. Zhang, “Solving large scale linear prediction problems using stochastic gradient descent algorithms,” in Proc. of the twenty-first int. conf. on Machine learning
2004
Cited alongside, same era.
M. Rásonyi, L. Stettner, et al
2005
Cited alongside, same era.
M. Pontil, Y. Ying, and D.-X. Zhou, “Error analysis for online gradient descent algorithms in reproducing kernel hilbert spaces,” tech. rep., Tech. Report, Dep. of Comp. Sci., Univ. College London, 2005
2005
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research
2013
Later among the works it cites.
M. P. Deisenroth, G. Neumann, J. Peters, et al
2013
Later among the works it cites.
2013
Later among the works it cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization
2013
Later among the works it cites.
G. Lever and R. Stafford, “Modelling policies in mdps in reproducing kernel hilbert space,” in A. I. and Statistics
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in neural information processing systems
2009
Cited alongside, same era.
S. Bhatnagar, D. Precup, D. Silver, R. S. Sutton, H. R. Maei, and C. Szepesvári, “Convergent temporal-difference learning with arbitrary smooth function approximation,” in Advances in Neural Information Processing Systems
2009
Cited alongside, same era.
A. Argyriou, C. A. Micchelli, and M. Pontil, “When is there a representer theorem? vector versus matrix regularizers,” Journal of Machine Learning Research
2009
Cited alongside, same era.
Cambridge university press, 2010
R. Durrett, Probability: theory and examples · 2010
Cited alongside, same era.
Y. Nesterov and V. Spokoiny, “Random gradient-free minimization of convex functions,” tech. rep., Université catholique de Louvain, Center for Operations Research and Econometrics (CORE), 2011
2011
Cited alongside, same era.
E. Tolstaya, A. Koppel, E. Stump, and A. Ribeiro, “Nonparametric stochastic compositional gradient descent for q-learning in continuous markov decision problems,”
Cited in the paper.
“Openai gym– continuous mountine car.” https://gym.openai.com/envs/MountainCarContinuous-v0/
Cited in the paper.
2016
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Nesterov and V. Spokoiny, “Random gradient-free minimization of convex functions,” Foundations of Computational Mathematics
2017
Later among the works it cites.
“Openai gym– cartpole.” https://gym.openai.com/envs/CartPole-v0/
2017
Later among the works it cites.