Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
Optimal oracle inequalities for solving projected fixed-point equations
Original
Wenlong Mou, Ashwin Pananjady, and Martin J. Wainwright · 2020
Later among the works it cites.
Least squares regression with markovian data: Fundamental limits and algorithms
Dheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain, and Praneeth Netrapalli · 2020
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Optimal policy evaluation using kernel-based temporal difference methods
Original
Yaqi Duan, Mengdi Wang, and Martin J Wainwright · 2021
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value functionaapproximation
Original
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Later among the works it cites.
Accelerated and instance-optimal policy evaluation with linear function approximation
Original
Tianjiao Li, Guanghui Lan, and Ashwin Pananjady · 2021
Later among the works it cites.
Asymptotically exact error characterization of offline policy evaluation with misspecified linear models
Kohei Miyaguchi · 2021
Later among the works it cites.
Optimal and instance-dependent guarantees for markovian linear stochastic approximation
Original
Wenlong Mou, Ashwin Pananjady, Martin J Wainwright, and Peter L Bartlett · 2021
Later among the works it cites.
Stabilizing dynamical systems via policy gradient methods
Juan Perdomo, Jack Umenberger, and Max Simchowitz · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch RL can be exponentially harder than online RL
Andrea Zanette · 2021
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Original
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason D Lee · 2022
Closest in time.