Fetching the paper…
Reading the bibliography…
Most of the existing works for reinforcement learning (RL) with general function approximation (FA) focus on understanding the statistical complexity or regret bounds.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2006
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin F Yang, and Ruslan Salakhutdinov · 2006
Earlier work this paper cites.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2008
Earlier work this paper cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2009
Earlier work this paper cites.
Is plug-in solver sample-efficient for feature-based reinforcement learning?
Qiwen Cui and Lin F Yang · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Universal ε \varepsilon -approximators for integrals
Michael Langberg and Leonard J Schulman · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Istvan Szita and Csaba Szepesvari · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Minimax sample complexity for turn-based stochastic game
Qiwen Cui and Lin F Yang · 2011
Earlier work this paper cites.
A unified framework for approximating and clustering data
Dan Feldman and Michael Langberg · 2011
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Turning big data into tiny data: constant-size coresets for k-means, pca and projective clustering
Dan Feldman, Melanie Schmidt, and Christian Sohler · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Concurrent pac rl
Zhaohan Guo and Emma Brunskill · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Michael B Cohen, Cameron Musco, and Jakub Pachocki · 2016
Earlier work this paper cites.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Active learning for cost-sensitive classification
Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daumé III, and John Langford · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
What doubling tricks can and can’t do for multi-armed bandits
Lilian Besson and Emilie Kaufmann · 2018
Cited alongside, same era.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin Yang, and Yinyu Ye · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dylan J Foster, Alekh Agarwal, Miroslav Dudik, Luo Haipeng, and Robert E Schapire · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2019
Cited alongside, same era.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Cited alongside, same era.
Xiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li, and Liwei Wang · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in rl
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Provably correct optimization and exploration with non-linear policies
Fei Feng, Wotao Yin, Alekh Agarwal, and Lin F Yang · 2021
Closest in time.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Closest in time.
A provably efficient algorithm for linear markov decision process with low switching cost
Minbo Gao, Tianle Xie, Simon S Du, and Lin F Yang · 2021
Closest in time.
Towards general function approximation in zero-sum markov games
Baihe Huang, Jason D Lee, Zhaoran Wang, and Zhuoran Yang · 2021
Closest in time.
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Gen Li, Laixi Shi, Yuxin Chen, Yuantao Gu, and Yuejie Chi · 2021
Closest in time.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Closest in time.
Ucb momentum q-learning: Correcting the bias without forgetting
Pierre Ménard, Omar Darwiche Domingues, Xuedong Shang, and Michal Valko · 2021
Closest in time.
On reward-free rl with kernel and neural function approximations: Single-agent mdp and markov game
Shuang Qiu, Jieping Ye, Zhaoran Wang, and Zhuoran Yang · 2021
Closest in time.
Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.
Randomized exploration is near-optimal for tabular mdp
Zhihan Xiong, Ruoqi Shen, and Simon S Du · 2021
Closest in time.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Closest in time.
Nearly minimax optimal reinforcement learning with linear function approximation
Pihe Hu, Yu Chen, and Longbo Huang · 2022
Closest in time.
Towards deployment-efficient reinforcement learning: Lower bound and optimality
Jiawei Huang, Jinglin Chen, Li Zhao, Tao Qin, Nan Jiang, and Tie-Yan Liu · 2022
Closest in time.
The power of exploiter: Provable multi-agent rl in large state spaces
Chi Jin, Qinghua Liu, and Tiancheng Yu · 2022
Closest in time.
Sample-efficient reinforcement learning with loglog (t) switching cost
Dan Qiao, Ming Yin, Ming Min, and Yu-Xiang Wang · 2022
Closest in time.
Reward-free rl is no harder than reward-aware rl in linear markov decision processes
Andrew J Wagenmaker, Yifang Chen, Max Simchowitz, Simon Du, and Kevin Jamieson · 2022
Closest in time.