Fetching the paper…
Reading the bibliography…
As representation learning becomes a powerful technique to reduce sample complexity in reinforcement learning (RL) in practice, theoretical understanding of its advantage is still limited.
Complexity analysis of real-time reinforcement learning
Koenig, S. and Simmons, R. G. (1993) · 1993
Earlier work this paper cites.
Few-shot learning via learning the representation, provably
Du, S. S., Hu, W., Kakade, S. M., Lee, J. D., and Lei, Q. (2020) · 2002
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L. (2006) · 2006
Earlier work this paper cites.
Provably efficient reward-agnostic navigation with linear value iteration
Zanette, A., Lazaric, A., Kochenderfer, M. J., and Brunskill, E. (2020b) · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Speedy q-learning
Azar, M. G., Munos, R., Ghavamzadeh, M., and Kappen, H. J. (2011) · 2011
Earlier work this paper cites.
PAC bounds for discounted mdps
Lattimore, T. and Hutter, M. (2012) · 2012
Earlier work this paper cites.
Sample complexity of multi-task reinforcement learning
Brunskill, E. and Li, L. (2013) · 2013
Earlier work this paper cites.
Sparse multi-task reinforcement learning
Calandriello, D., Lazaric, A., and Restelli, M. (2015) · 2015
Earlier work this paper cites.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B. (2016) · 2016
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Osband, I., Roy, B. V., and Wen, Z. (2016) · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Earlier work this paper cites.
Multi-task learning for contextual bandits
Deshmukh, A. A., Dogan, Ü., and Scott, C. (2017) · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Teh, Y. W., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017) · 2017
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Earlier work this paper cites.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Sidford, A., Wang, M., Wu, X., and Ye, Y. (2018) · 2018
Earlier work this paper cites.
Provably efficient rl with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J. (2019) · 2019
Earlier work this paper cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J. (2019a) · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W. (2020) · 2020
Cited alongside, same era.
Provable representation learning for imitation learning via bi-level optimization
Arora, S., Du, S. S., Kakade, S. M., Luo, Y., and Saunshi, N. (2020) · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2020) · 2020
Cited alongside, same era.
Sharing knowledge in multi-task deep reinforcement learning
D’Eramo, C., Tateo, D., Bonarini, A., Restelli, M., and Peters, J. (2020) · 2020
Cited alongside, same era.
Representation learning for online and offline rl in low-rank mdps
Uehara, M., Zhang, X., and Sun, W. (2021) · 2021
Later among the works it cites.
What are the statistical limits of offline RL with linear function approximation?
Wang, R., Foster, D. P., and Kakade, S. M. (2021a) · 2021
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A. (2021b) · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C., Jiang, N., Mineiro, P., and Agarwal, A. (2021a) · 2021
Later among the works it cites.
Impact of representation learning in linear bandits
Yang, J., Hu, W., Lee, J. D., and Du, S. S. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2020) · 2020
Cited alongside, same era.
Meta-learning for mixed linear regression
Kong, W., Somani, R., Song, Z., Kakade, S. M., and Oh, S. (2020) · 2020
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, D., Henaff, M., Krishnamurthy, A., and Langford, J. (2020) · 2020
Cited alongside, same era.
A unifying view of optimism in episodic reinforcement learning
Neu, G. and Pike-Burke, C. (2020) · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. and Wang, M. (2020) · 2020
Cited alongside, same era.
Multi-task and meta-learning with sparse linear bandits
Cella, L. and Pontil, M. (2021) · 2021
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R. (2021) · 2021
Cited alongside, same era.
Optimal uniform OPE and model-based offline reinforcement learning in time-homogeneous, reward-free and task-agnostic settings
Yin, M. and Wang, Y. (2021a) · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch RL can be exponentially harder than online RL
Zanette, A. (2021) · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E. (2021) · 2021
Later among the works it cites.
Provably efficient representation learning in low-rank markov decision processes
Zhang, W., He, J., Zhou, D., Zhang, A., and Gu, Q. (2021) · 2021
Later among the works it cites.
Provable benefits of representational transfer in reinforcement learning
Agarwal, A., Song, Y., Sun, W., Wang, K., Wang, M., and Zhang, X. (2022) · 2022
Closest in time.
All you need is supervised learning: From imitation learning to meta-rl with upside down RL
Arulkumaran, K., Ashley, D. R., Schmidhuber, J., and Srivastava, R. K. (2022) · 2022
Closest in time.
Non-stationary bandits and meta-learning with a small set of optimal arms
Azizi, M. J., Duong, T., Abbasi-Yadkori, Y., György, A., Vernade, C., and Ghavamzadeh, M. (2022) · 2022
Closest in time.
Multi-task representation learning with stochastic linear bandits
Cella, L., Lounici, K., and Pontil, M. (2022) · 2022
Closest in time.
Provable general function class representation learning in multitask bandits and mdps
Lu, R., Zhao, A., Du, S. S., and Huang, G. (2022) · 2022
Closest in time.
Meta learning mdps with linear transition models
Müller, R. and Pacchiano, A. (2022) · 2022
Closest in time.
Non-stationary representation learning in sequential linear bandits
Qin, Y., Menara, T., Oymak, S., Ching, S., and Pasqualetti, F. (2022) · 2022
Closest in time.
Yin, M., Duan, Y., Wang, M., and Wang, Y. (2022) · 2022
Closest in time.
Efficient reinforcement learning in block mdps: A model-free representation learning approach
Zhang, X., Song, Y., Uehara, M., Wang, M., Sun, W., and Agarwal, A. (2022) · 2022
Closest in time.