Fetching the paper…
Reading the bibliography…
Policy evaluation or value function or Q-function approximation is a key procedure in reinforcement learning (RL).
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Generalized polynomial approximations in markovian decision processes
P. J. Schweitzer and A. Seidmann · 1985
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
A generalized representer theorem
B. Schölkopf, R. Herbrich, and A. Smola · 2001
Earlier work this paper cites.
Classes of kernels for machine learning: a statistics perspective
M. G. Genton · 2001
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
M. Belkin and P. Niyogi · 2003
Earlier work this paper cites.
Kernels and regularization on graphs
A. J. Smola and R. Kondor · 2003
Earlier work this paper cites.
Beyond the point cloud: from transductive to semi-supervised learning
V. Sindhwani, P. Niyogi, and M. Belkin · 2005
Earlier work this paper cites.
Manifold regularization: A geometric framework for learning from labeled and unlabeled examples
M. Belkin, P. Niyogi, and V. Sindhwani · 2006
Earlier work this paper cites.
Kernel-based least squares policy iteration for reinforcement learning
X. Xu, D. Hu, and X. Lu · 2007
Cited alongside, same era.
An analysis of laplacian methods for value function approximation in mdps
M. Petrik · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Cited alongside, same era.
Geodesic gaussian kernels for value function approximation
M. Sugiyama, H. Hachiya, C. Towell, and S. Vijayakumar · 2008
Cited alongside, same era.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
R. Parr, L. Li, G. Taylor, C. Painter-Wakefield, and M. L. Littman · 2008
Cited alongside, same era.
Regularized policy iteration
Representation policy iteration
S. Mahadevan · 2012
Later among the works it cites.
Regularized least squares temporal difference learning with nested l2 and l1 penalization
M. Hoffman, A. Lazaric, M. Ghavamzadeh, and R. Munos · 2012
Later among the works it cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
Optimization and stabilization of trajectories for constrained dynamical systems
M. Posa, S. Kuindersma, and R. Tedrake · 2016
Later among the works it cites.
Orthogonal random features
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. M. Farahmand, M. Ghavamzadeh, S. Mannor, and C. Szepesvári · 2009
Cited alongside, same era.
Kernelized value function approximation for reinforcement learning
G. Taylor and R. Parr · 2009
Cited alongside, same era.
Regularization and feature selection in least-squares temporal difference learning
J. Z. Kolter and A. Y. Ng · 2009
Cited alongside, same era.
Semi-supervised learning by higher order regularization
X. Zhou and M. Belkin · 2011
Cited alongside, same era.
X. Y. Felix, A. T. Suresh, K. M. Choromanski, D. N. Holtmann-Rice, and S. Kumar · 2016
Later among the works it cites.
Learning from conditional distributions via dual kernel embeddings
B. Dai, N. He, Y. Pan, B. Boots, and L. Song · 2016
Later among the works it cites.
Manifold regularized reinforcement learning
H. Li, D. Liu, and D. Wang · 2017
Closest in time.
The unreasonable effectiveness of random orthogonal embeddings
K. Choromanski, M. Rowland, and A. Weller · 2017
Closest in time.