Fetching the paper…
Reading the bibliography…
Reward-free reinforcement learning (RL) considers the setting where the agent does not have access to a reward function during exploration, but must propose a near-optimal policy for an arbitrary reward function revealed only after exploring.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A., Brandfonbrener, D., Brunskill, E., Pirotta, M., and Lazaric, A · 1964
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A · 1975
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Provably efficient reward-agnostic navigation with linear value iteration
Zanette, A., Lazaric, A., Kochenderfer, M. J., and Brunskill, E · 2008
Earlier work this paper cites.
Introduction to nonparametric estimation., 2009
Tsybakov, A. B · 2009
Earlier work this paper cites.
Zhang, Z., Ji, X., and Du, S. S · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Nearly minimax optimal reward-free reinforcement learning
Zhang, Z., Du, S. S., and Ji, X · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
On the complexity of bandit and derivative-free stochastic convex optimization
Shamir, O · 2013
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2019
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 2019
Cited alongside, same era.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A · 2019
Cited alongside, same era.
On reward-free reinforcement learning with linear function approximation
Wang, R., Du, S. S., Yang, L. F., and Salakhutdinov, R · 2020
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D., Gu, Q., and Szepesvari, C · 2020
Later among the works it cites.
Agarwal, N., Chaudhuri, S., Jain, P., Nagaraj, D., and Netrapalli, P · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Jia, Z., Yang, L., Szepesvari, C., and Wang, M · 2020
Cited alongside, same era.
Fast active learning for pure exploration in reinforcement learning
Ménard, P., Domingues, O. D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M · 2020
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S · 2020
Cited alongside, same era.
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A · 2021
Later among the works it cites.
Online sparse reinforcement learning
Hao, B., Lattimore, T., Szepesvári, C., and Wang, M · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Modi, A., Chen, J., Krishnamurthy, A., Jiang, N., and Agarwal, A · 2021
Later among the works it cites.
An exponential lower bound for linearly-realizable mdps with constant suboptimality gap
Wang, Y., Wang, R., and Kakade, S. M · 2021
Later among the works it cites.
Exponential lower bounds for planning in mdps with linearly-realizable optimal action-value functions
Weisz, G., Amortila, P., and Szepesvári, C · 2021
Later among the works it cites.
Gap-dependent unsupervised exploration for reinforcement learning
Wu, J., Braverman, V., and Yang, L. F · 2021
Later among the works it cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D., He, J., and Gu, Q · 2021
Later among the works it cites.
Towards deployment-efficient reinforcement learning: Lower bound and optimality
Huang, J., Chen, J., Zhao, L., Qin, T., Jiang, N., and Liu, T.-Y · 2022
Closest in time.