Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms based on high-dimensional function approximation have achieved tremendous empirical success in large-scale problems with an enormous number of states.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Yuhong Yang and Andrew Barron · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
The covering number in learning theory
Ding-Xuan Zhou · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Capacity of reproducing kernel spaces in learning theory
Ding-Xuan Zhou · 2003
Earlier work this paper cites.
Local rademacher complexities
Peter L Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Q-learning with linear function approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
Multivariate L ∞ {L}_{\infty} approximation in the worst case setting over reproducing kernel Hilbert spaces
Frances Y Kuo, Grzegorz W Wasilkowski, and Henryk Woźniakowski · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Generalization bounds for learning kernels
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2010
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J. Kappen · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck · 2012
Earlier work this paper cites.
Learning kernels using local rademacher complexity
Corinna Cortes, Marius Kloft, and Mehryar Mohri · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Later among the works it cites.
Barron spaces and the compositional function spaces for neural network models
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
A priori estimates of the population risk for two-layer neural networks
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Analysis of classification-based policy iteration algorithms
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Cited alongside, same era.
An introduction to the theory of reproducing kernel Hilbert spaces
Vern I Paulsen and Mrinal Raghupathi · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Later among the works it cites.
Deep neural tangent kernel and laplace kernel have the same RKHS
Lin Chen and Sheng Xu · 2020
Later among the works it cites.
Regret bounds for kernel-based reinforcement learning
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2020
Later among the works it cites.
Weinan E, Chao Ma, Stephan Wojtowytsch, and Lei Wu · 2020
Later among the works it cites.
A theoretical analysis of deep Q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Later among the works it cites.
On the similarity between the laplace and neural tangent kernels
Amnon Geifman, Abhay Yadav, Yoni Kasten, Meirav Galun, David Jacobs, and Ronen Basri · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Provably efficient reinforcement learning with general value function approximation
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with kernel and neural function approximations
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael Jordan · 2020
Later among the works it cites.
On function appproximation in reinforcement learning: Optimisim in the face of large state spaces
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Kolmogorov width decay and poor approximators in machine learning: Shallow neural networks, random feature models and neural tangent kernels
Weinan E and Stephan Wojtowytsch · 2021
Closest in time.