Fetching the paper…
Reading the bibliography…
We study reinforcement learning (RL) with linear function approximation.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A., Brandfonbrener, D., Brunskill, E., Pirotta, M., and Lazaric, A · 1964
Earlier work this paper cites.
Best linear unbiased estimation and prediction under a selection model
Henderson, C. R · 1975
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Linear multi-resource allocation with semi-bandit feedback
Lattimore, T., Crammer, K., and Szepesvári, C · 2015
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Earlier work this paper cites.
On oracle-efficient pac rl with rich observations
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, J. and Krause, A · 2018
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S · 2020
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. and Wang, M · 2020
Later among the works it cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z., Zhou, Y., and Ji, X · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Wang, T., Zhou, D., and Gu, Q · 2021
Later among the works it cites.
Vo q q l: Towards optimal regret in model-free rl with nonlinear function approximation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E · 2019
Cited alongside, same era.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Jia, Z., Yang, L., Szepesvari, C., and Wang, M · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
He, J., Zhou, D., and Gu, Q
Cited in the paper.
Nearly minimax optimal reinforcement learning for discounted mdps
He, J., Zhou, D., and Gu, Q
Cited in the paper.
Agarwal, A., Jin, Y., and Zhang, T · 2022
Closest in time.
Nearly optimal algorithms for linear contextual bandits with adversarial corruptions
He, J., Zhou, D., Zhang, T., and Gu, Q · 2022
Closest in time.
Nearly minimax optimal reinforcement learning with linear function approximation
Hu, P., Chen, Y., and Huang, L · 2022
Closest in time.
Computationally efficient horizon-free reinforcement learning for linear mixture mdps
Zhou, D. and Gu, Q · 2022
Closest in time.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D., He, J., and Gu, Q · 2022
Closest in time.