Fetching the paper…
Reading the bibliography…
We consider the problem of offline reinforcement learning (RL) -- a well-motivated setting of RL that aims at policy optimization using only historical data.
Zanette, A., & Brunskill, E. (2019) · 1901
Earlier work this paper cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., & Yue, Y. (2019) · 1903
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J., & Jiang, N. (2019) · 1905
Earlier work this paper cites.
On the optimality of sparse model-based planning for markov decision processes
Agarwal, A., Kakade, S., & Yang, L. F. (2019) · 1906
Earlier work this paper cites.
Variance-reduced q q -learning is minimax optimal
Wainwright, M. J. (2019) · 1906
Earlier work this paper cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N., & Uehara, M. (2019a) · 1908
Earlier work this paper cites.
Kallus, N., & Uehara, M. (2019b) · 1909
Earlier work this paper cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., & Jiang, N. (2019) · 1910
Earlier work this paper cites.
Learning with good feature representations in bandits and in rl with a generative model
Lattimore, T., & Szepesvari, C. (2019) · 1911
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., & Wang, Z. (2019) · 1912
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Chernoff, H., et al. (1952) · 1952
Earlier work this paper cites.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., & Wang, M. (2020) · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R. (2003) · 2003
Earlier work this paper cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T., & Jiang, N. (2020b) · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z., Zhou, Y., & Ji, X. (2020) · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020) · 2005
Earlier work this paper cites.
Concentration inequalities and martingale inequalities: a survey
Chung, F., & Lu, L. (2006) · 2006
Earlier work this paper cites.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
Li, G., Wei, Y., Chi, Y., Gu, Y., & Chen, Y. (2020) · 2006
Cited alongside, same era.
Bias in error estimation when using cross-validation for model selection
Varma, S., & Simon, R. (2006) · 2006
Cited alongside, same era.
Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., & Brunskill, E. (2020b) · 2007
Cited alongside, same era.
Batch value-function approximation with only realizability
Xie, T., & Jiang, N. (2020a) · 2008
Cited alongside, same era.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., & Munos, R. (2017) · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., & Brunskill, E. (2017) · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., & Zhou, D. (2017) · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., & Schapire, R. E. (2017) · 2017
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., & Brunskill, E. (2018) · 2018
Later among the works it cites.
Is q-learning provably efficient?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Domingues, O. D., Ménard, P., Kaufmann, E., & Valko, M. (2020) · 2010
Cited alongside, same era.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D. P., & Kakade, S. M. (2020) · 2010
Cited alongside, same era.
Freedman’s inequality for matrix martingales
Tropp, J., et al. (2011) · 2011
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., & Wang, Z. (2020) · 2012
Cited alongside, same era.
Batch reinforcement learning
Lange, S., Gabel, T., & Riedmiller, M. (2012) · 2012
Cited alongside, same era.
Zanette, A. (2020) · 2012
Cited alongside, same era.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., & Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R., & Zhang, T. (2013) · 2013
Cited alongside, same era.
Jin, C., Allen-Zhu, Z., Bubeck, S., & Jordan, M. I. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., & Zhou, D. (2018) · 2018
Later among the works it cites.
Provably efficient q-learning with low switching cost
Bai, Y., Xie, T., Jiang, N., & Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Tight regret bounds for model-based reinforcement learning with greedy policies
Efroni, Y., Merlis, N., Ghavamzadeh, M., & Mannor, S. (2019) · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., & Brunskill, E. (2019) · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M., & Jamieson, K. G. (2019) · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., & Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Yang, L., & Wang, M. (2019) · 2019
Later among the works it cites.
Accountable off-policy evaluation with kernel bellman statistics
Feng, Y., Ren, T., Tang, Z., & Liu, Q. (2020) · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Sidford, A., Wang, M., Yang, L., & Ye, Y. (2020) · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M., & Wang, Y.-X. (2020) · 2020
Later among the works it cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
Yin, M., Bai, Y., & Wang, Y.-X. (2021) · 2021
Closest in time.