Fetching the paper…
Reading the bibliography…
We consider policy evaluation in infinite-horizon discounted Markov decision problems (MDPs) with infinite spaces.
H. Robbins and S. Monro, “A stochastic approximation method,” Ann. Math. Statist. , vol. 22, no. 3, pp. 400–407, 09 1951
1951
Earlier work this paper cites.
R. Bellman, Dynamic Programming , 1st ed. Princeton, NJ, USA: Princeton University Press, 1957
1957
Earlier work this paper cites.
G. Kimeldorf and G. Wahba, “Some results on tchebycheffian spline functions,” J Math Anal Appl , vol. 33, no. 1, pp. 82–95, 1971
1971
Earlier work this paper cites.
R. Wheeden, R. Wheeden, and A. Zygmund, Measure and Integral: An Introduction to Real Analysis , ser. Chapman & Hall/CRC Pure and Applied Mathematics. Taylor & Francis, 1977. [Online]. Available: https://books.google.com/books?id=YDkDmQ_hdmcC
1977
Earlier work this paper cites.
D. P. Bertsekas and S. E. Shreve, Stochastic optimal control: The discrete time case . Academic Press, 1978, vol. 23
1978
Earlier work this paper cites.
Y. Ermoliev, “Stochastic quasigradient methods and their application to system optimization,” Stochastics , vol. 9, no. 1-2, pp. 1–36, 1983
1983
Earlier work this paper cites.
A. Korostelev, “Stochastic recurrent procedures: Local properties,” Nauka: Moscow (in Russian) , 1984
1984
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine learning , vol. 3, no. 1, pp. 9–44, 1988
1988
Earlier work this paper cites.
E. Rimon and D. E. Koditschek, “Exact robot navigation using artificial potential functions,” Departmental Papers (ESE) , p. 323, 1992
1992
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SICON , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
Y. Pati, R. Rezaiifar, and P. Krishnaprasad, “Orthogonal Matching Pursuit: Recursive Function Approximation with Applications to Wavelet Decomposition,” in Asilomar Conference , 1993
1993
Earlier work this paper cites.
J. N. Tsitsiklis, “Asynchronous stochastic approximation and q-learning,” Machine Learning , vol. 16, no. 3, pp. 185–202, 1994
1994
Earlier work this paper cites.
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in ICML . Morgan Kaufmann, 1995, pp. 30–37
1995
Earlier work this paper cites.
S. J. Bradtke and A. G. Barto, “Linear least-squares algorithms for td learning,” Machine learning , vol. 22, no. 1-3, pp. 33–57, 1996
1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “An analysis of temporal-difference learning with function approximation,” IEEE Trans. Autom. Control , vol. 42, no. 5, pp. 674–690, 1997
1997
Earlier work this paper cites.
V. S. Borkar and V. R. Konda, “The actor-critic algorithm as multi-time-scale stochastic approximation,” Sadhana , vol. 22, no. 4, pp. 525–543, 1997
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press Cambridge, 1998, vol. 1, no. 1
1998
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML , 1999, pp. 278–287
1999
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in NeurIPS , 2000, pp. 1057–1063
2000
Earlier work this paper cites.
D. Precup, “Eligibility traces for off-policy policy evaluation,” Computer Science Department Faculty Publication Series , p. 80, 2000
2000
Cited alongside, same era.
V. S. Borkar and S. P. Meyn, “The ode method for convergence of stochastic approximation and reinforcement learning,” SICON , vol. 38, no. 2, pp. 447–469, 2000
2000
Cited alongside, same era.
D. Ormoneit and Ś. Sen, “Kernel-based reinforcement learning,” Machine learning , vol. 49, no. 2-3, pp. 161–178, 2002
2002
Cited alongside, same era.
D.-X. Zhou, “The covering number in learning theory,” Complexity , vol. 18, no. 3, pp. 739–767, 2002
2002
Cited alongside, same era.
P. Vincent and Y. Bengio, “Kernel matching pursuit,” Machine Learning , vol. 48, no. 1, pp. 165–187, 2002
2002
Cited alongside, same era.
V. Norkin and M. Keyzer, “On stochastic optimization and statistical learning in reproducing kernel hilbert spaces by support vector machines (svm),” Informatica , vol. 20, no. 2, pp. 273–292, 2009
2009
Later among the works it cites.
H. Brezis, Functional analysis, Sobolev spaces and partial differential equations . Springer Science & Business Media, 2010
2010
Later among the works it cites.
W. B. Powell and J. Ma, “A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications,” J. of Control Theory & \& Apps. , vol. 9, no. 3, pp. 336–352, 2011
2011
Later among the works it cites.
S. Grünewälder, G. Lever, L. Baldassarre, M. Pontil, and A. Gretton, “Modelling transition dynamics in mdps with rkhs embeddings,” in ICML , vol. 1, 2012, pp. 535–542
2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Engel, S. Mannor, and R. Meir, “Bayes meets bellman: The gaussian process approach to temporal difference learning,” in ICML , 2003
2003
Cited alongside, same era.
V. R. Konda and J. N. Tsitsiklis, “On actor-critic algorithms,” SIAM J. Control & \& Optimization , vol. 42, no. 4, pp. 1143–1166, 2003
2003
Cited alongside, same era.
J. Kivinen, A. J. Smola, and R. C. Williamson, “Online Learning with Kernels,” IEEE Trans. Signal Process. , vol. 52, pp. 2165–2176, August 2004
2004
Cited alongside, same era.
Y. Engel, S. Mannor, and R. Meir, “The kernel recursive least-squares algorithm,” IEEE Trans. Signal Process. , vol. 52, no. 8, pp. 2275–2285, Aug 2004
2004
Cited alongside, same era.
C. A. Micchelli, Y. Xu, and H. Zhang, “Universal kernels,” JMLR , vol. 7, no. Dec, pp. 2651–2667, 2006
2006
Cited alongside, same era.
N. K. Jong and P. Stone, “Model-based function approximation in reinforcement learning,” in AAMAS . ACM, 2007, p. 95
2007
Cited alongside, same era.
A. Smola, A. Gretton, L. Song, and B. Schölkopf, “A hilbert space embedding for distributions,” in International Conference on Algorithmic Learning Theory . Springer, 2007, pp. 13–31
2007
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” Int. J. Robotics Res. , p. 0278364913495721, 2013
2013
Later among the works it cites.
2014
Later among the works it cites.
A. Shapiro and D. Dentcheva, Lectures on stochastic programming: modeling and theory . Siam, 2014, vol. 16
2014
Later among the works it cites.
M. Wang and D. P. Bertsekas, “Incremental constraint projection-proximal methods for nonsmooth convex optimization,” SIOPT , 2014
2014
Later among the works it cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in ICML , 2015, pp. 1889–1897
2015
Later among the works it cites.
A.-m. Farahmand, C. Ghavamzadeh, Mohammadand Szepesvári, and S. Mannor, “Regularized policy iteration with nonparametric function spaces,” JMLR , vol. 17, no. 139, pp. 1–66, 2016
2016
Later among the works it cites.
G. Lever, J. Shawe-Taylor, R. Stafford, and C. Szepesvari, “Compressed conditional mean embeddings for model-based reinforcement learning,” in AAAI , 2016
2016
Later among the works it cites.
M. Wang, J. Liu, and E. Fang, “Accelerating stochastic composition optimization,” in NeurIPS , 2016, pp. 1714–1722
2016
Later among the works it cites.
B. Dai, N. He, Y. Pan, B. Boots, and L. Song, “Learning from conditional distributions via dual embeddings,” in Artificial Intelligence and Statistics , 2017, pp. 1458–1467
2017
Closest in time.
M. Wang, E. X. Fang, and H. Liu, “Stochastic compositional gradient descent: Algorithms for minimizing compositions of expected-value functions,” Math Program , vol. 161, no. 1-2, pp. 419–449, 2017
2017
Closest in time.
E. Tolstaya, A. Koppel, E. Stump, and A. Ribeiro, “Nonparametric stochastic compositional gradient descent for q-learning in continuous markov decision problems,” in 2018 Annual American Control Conference (ACC) . IEEE, 2018, pp. 6608–6615
2018
Closest in time.
A. Koppel, G. Warnell, E. Stump, and A. Ribeiro, “Parsimonious online learning with kernels via sparse projections in function space,” JMLR , vol. 20, no. 1, pp. 83–126, 2019
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
S. Bhatt, A. Koppel, and V. Krishnamurthy, “Policy gradient using weak derivatives for reinforcement learning,” in 2019 53rd Annual Conference on Information Sciences and Systems (CISS) , March 2019, pp. 1–3
2019
Closest in time.