Fetching the paper…
Reading the bibliography…
Value function learning plays a central role in many state-of-the-art reinforcement-learning algorithms.
Positive definite functions and generalizations, an historical survey
Stewart, J · 1976
Earlier work this paper cites.
On U-statistics and v. mise’ statistics for weakly dependent processes
Denker, M. and Keller, G · 1983
Earlier work this paper cites.
Stochastic Systems: Estimation, Identification, and Adaptive Control
Kumar, P. and Varaiya, P · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. C · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. A. and Moore, A. W · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, G. J · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B · 1997
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A · 1999
Earlier work this paper cites.
Learning with kernels: Support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2001
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś · 2002
Earlier work this paper cites.
Kernel least-squares temporal difference learning
Xu, X., Xie, T., Hu, D., and Lu, X · 2005
Earlier work this paper cites.
Kernel-based least-squares policy iteration for reinforcement learning
Xu, X., Hu, D., and Lu, X · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimizing based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Regularized policy iteration
Farahmand, A. M., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2008
Earlier work this paper cites.
Finite-time bounds for sampling-based fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Maei, H. R., Szepesvári, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. S · 2009
Cited alongside, same era.
Approximation theorems of mathematical statistics , volume 162
Serfling, R. J · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H., Precup, D., Bhatnagar, S., Szepesvári, C., and Wiewiora, E · 2009
Cited alongside, same era.
Kernelized value function approximation for reinforcement learning
Taylor, G. and Parr, R · 2009
Cited alongside, same era.
Nonlinear Programming
Bertsekas, D. P · 2016
Later among the works it cites.
A kernel test of goodness of fit
Chwialkowski, K., Strathmann, H., and Gretton, A · 2016
Later among the works it cites.
Regularized policy iteration with nonparametric function spaces
Farahmand, A. M., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2016
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Later among the works it cites.
Continuous deep Q-learning with model-based acceleration
Gu, S., Lillicrap, T. P., Sutskever, I., and Levine, S · 2016
Later among the works it cites.
A kernelized Stein discrepancy for goodness-of-fit tests
Liu, Q., Lee, J., and Jordan, M · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maei, H. R., Szepesvári, C., Bhatnagar, S., and Sutton, R. S · 2010
Cited alongside, same era.
Hilbert space embeddings and metrics on probability measures
Sriperumbudur, B. K., Gretton, A., Fukumizu, K., Schölkopf, B., and Lanckriet, G. R · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvári, C · 2010
Cited alongside, same era.
Reproducing kernel Hilbert spaces in probability and statistics
Berlinet, A. and Thomas-Agnan, C · 2011
Cited alongside, same era.
Regularized least squares temporal difference learning with nested ℓ 2 \ell_{2} and ℓ 1 \ell_{1} penalization
Hoffman, M. W., Lazaric, A., Ghavamzadeh, M., and Munos, R · 2011
Cited alongside, same era.
Gradient Temporal-Difference Learning Algorithms
Maei, H. R · 2011
Cited alongside, same era.
Deriving the asymptotic distribution of U- and V-statistics of dependent data using weighted empirical processes
Beutner, E. and Zähle, H · 2012
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Adrià, Badia, P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Dai, B., He, N., Pan, Y., Boots, B., and Song, L · 2017
Later among the works it cites.
Kernel mean embedding of distributions: A review and beyond
Muandet, K., Fukumizu, K., Sriperumbudur, B., Schölkopf, B., et al · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
Wang, M · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Wu, Y., Mansimov, E., Grosse, R. B., Liao, S., and Ba, J · 2017
Later among the works it cites.
Scalable bilinear π \pi -learning using state and action features
Chen, Y., Li, L., and Wang, M · 2018
Later among the works it cites.
Trust-PCL: An off-policy trust region method for continuous control
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Closest in time.