Near optimal behavior via approximate state abstraction
Abel, D., Hershkowitz, D. E., and Littman, M. L · 2016
Later among the works it cites.
Regularized policy iteration with nonparametric function spaces
Farahmand, A.-m., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2016
Later among the works it cites.
The malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Later among the works it cites.
PAC reinforcement learning with rich observations
Krishnamurthy, A., Agarwal, A., and Langford, J · 2016
Later among the works it cites.
Policy error bounds for model-based reinforcement learning with factored linear models
Pires, B. Á. and Szepesvári, C · 2016
Later among the works it cites.
Contextual Decision Processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Later among the works it cites.
Boosted fitted q-iteration
Tosatto, S., Pirotta, M., D’Eramo, C., and Restelli, M · 2017
Later among the works it cites.
IEOR 8100: Reinforcement Learning. Lecture 4: Approximate Dynamic Programming
Agrawal, S · 2018
Later among the works it cites.
Scalable bilinear π \pi learning using state and action features
Original
Chen, Y., Li, L., and Wang, M · 2018
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L · 2018
Later among the works it cites.
On Oracle-Efficient PAC RL with Rich Observations
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2018
Later among the works it cites.
CS 598: Notes on State Abstractions
Jiang, N · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Later among the works it cites.
Non-delusional q-learning and value-iteration
Lu, T., Schuurmans, D., and Boutilier, C · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
Original
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Later among the works it cites.
Model-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J · 2019
Closest in time.
A Theoretical Analysis of Deep Q-Learning
Original
Yang, Z., Xie, Y., and Wang, Z · 2019
Closest in time.