Fetching the paper…
Reading the bibliography…
Direct policy gradient methods for reinforcement learning are a successful approach for a variety of reasons: they are model free, they directly optimize the performance metric of interest, and they allow for richly parameterized policies.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 1906
Earlier work this paper cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 1906
Earlier work this paper cites.
Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvari, and Gellert Weisz · 1908
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Policy search by dynamic programming
J. A. Bagnell, Sham M Kakade, Jeff G. Schneider, and Andrew Y. Ng · 2004
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang, and Lin F. Yang · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Online Markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun · 2009
Earlier work this paper cites.
A unifying framework for computational reinforcement learning theory
Lihong Li · 2009
Cited alongside, same era.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Gaussian process optimization in the bandit setting: no regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
On polynomial time PAC reinforcement learning with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Non-delusional q-learning and value-iteration
Tyler Lu, Dale Schuurmans, and Craig Boutilier · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
István Szita and Csaba Szepesvári · 2010
Cited alongside, same era.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Cited alongside, same era.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J. Kappen · 2012
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Michal Valko, Nathaniel Korda, Rémi Munos, Ilias Flaounas, and Nelo Cristianini · 2013
Cited alongside, same era.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Cited alongside, same era.
Modularized implementation of deep rl algorithms in pytorch
Zhang Shangtong · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift, 2019
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Provably efficient reinforcement learning with aggregated states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 2019
Later among the works it cites.
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Later among the works it cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Explicit explore-exploit algorithms in continuous state spaces
Mikael Henaff · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps, 2019
Lior Shani, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Closest in time.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Closest in time.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Closest in time.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Closest in time.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2020
Closest in time.