Mastering the game of Go with deep neural networks and tree search
Silver, D · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Original
Azar, M. G · 2017
Later among the works it cites.
On kernelized multi-armed bandits
Original
Chowdhury, S. R · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D · 2017
Later among the works it cites.
StarCraft II: A new challenge for reinforcement learning
Original
Vinyals, O · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q · 2018
Later among the works it cites.
Learning to optimize via information-directed sampling
Russo, D · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, R · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O · 2019
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Kumar, A · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Closest in time.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y · 2020
Closest in time.
A theoretical analysis of deep Q-learning
Fan, J · 2020
Closest in time.
Minimax value interval for off-policy evaluation and policy optimization
Jiang, N · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Closest in time.
Bandit algorithms
Lattimore, T · 2020
Closest in time.
Scalability in perception for autonomous driving: Waymo open dataset
Sun, P · 2020
Closest in time.
Minimax weight and Q-function learning for off-policy evaluation
Uehara, M · 2020
Closest in time.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M · 2020
Closest in time.
Cautiously optimistic policy optimization and exploration with linear function approximation
Zanette, A · 2021
Closest in time.