Fetching the paper…
Reading the bibliography…
In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model.
The Theory of Probabilities
Sergei Natanovich Bernstein · 1946
Earlier work this paper cites.
Probability Inequalities for Sums of Bounded Random Variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Kazuoki Azuma · 1967
Earlier work this paper cites.
Q-Learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On the Generation of Markov Decision Processes
TW Archibald, KIM McKinnon, and LC Thomas · 1995
Earlier work this paper cites.
Learning Rates for Q-learning
Eyal Even-Dar, Yishay Mansour, and Peter Bartlett · 2003
Earlier work this paper cites.
Speedy Q-Learning
Mohammad Azar, Mohammad Ghavamzadeh, Hilbert Kappen, and Rémi Munos · 2011
Earlier work this paper cites.
PAC Bounds for Discounted MDPs
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes
Bruno Scherrer and Boris Lesner · 2012
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Azar, Rémi Munos, and Hilbert J. Kappen · 2013
Earlier work this paper cites.
Concentration Inequalities - A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
Taming the Noise in Reinforcement Learning via Soft Updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Minimax Regret Bounds for Reinforcement Learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
A unified view of entropy-regularized Markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Softmax exploration strategies for multiobjective reinforcement learning
Peter Vamplew, Richard Dazeley, and Cameron Foale · 2017
Cited alongside, same era.
Is Q-Learning Provably Efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I. Jordan · 2018
Variance-reduced Q Q -learning is minimax optimal
Martin J Wainwright · 2019
Later among the works it cites.
A regularized approach to sparse optimal policy in reinforcement learning
Wenhao Yang, Xiang Li, and Zhihua Zhang · 2019
Later among the works it cites.
Model-Based Reinforcement Learning with a Generative Model is Minimax Optimal
Alekh Agarwal, Sham Kakade, and Lin F. Yang · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2020
Later among the works it cites.
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
On the Global Convergence Rates of Softmax Policy Gradient Methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse Markov decision processes with causal sparse tsallis entropy regularization for reinforcement learning
Kyungjae Lee, Sungjoon Choi, and Songhwai Oh · 2018
Cited alongside, same era.
Near-Optimal Time and Sample Complexities for Solving Markov Decision Processes with a Generative Model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
A Theory of Regularized Markov Decision Processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Cited alongside, same era.
Theoretical Analysis of Efficiency and Robustness of Softmax and Gap-Increasing Operators in Reinforcement Learning
Tadashi Kozuno, Eiji Uchibe, and Kenji Doya · 2019
Cited alongside, same era.
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
Gen Li, Changxiao Cai, Yuxin Chen, Yuantao Gu, Yuting Wei, and Yuejie Chi
Cited in the paper.
Polyak-Ruppert Averaged Q-Leaning is Statistically Efficient
Xiang Li, Wenhao Yang, Zhihua Zhang, and Michael I Jordan
Cited in the paper.
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning
Koulik Khamaru, Eric Xia, Martin J Wainwright, and Michael I Jordan · 2021
Later among the works it cites.
Nearly Minimax Optimal Reinforcement Learning for Linear Mixture Markov Decision Processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Closest in time.