Optimism in reinforcement learning with generalized linear function approximation
Original
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Original
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Later among the works it cites.
Minimax sample complexity for turn-based stochastic game
Original
Qiwen Cui and Lin F Yang · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Zeyu Jia, Lin Yang, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Original
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin Yang, and Yinyu Ye · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Original
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Original
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2020
Later among the works it cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2020
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.