Fetching the paper…
Reading the bibliography…
Classical theory in reinforcement learning (RL) predominantly focuses on the single task setting, where an agent learns to solve a task through trial-and-error experience, given access to data only from that task.
Markov decision processes
M. L. Puterman · 1990
Earlier work this paper cites.
Introduction: The challenge of reinforcement learning
R. S. Sutton · 1992
Earlier work this paper cites.
Learning internal representations
J. Baxter · 1995
Earlier work this paper cites.
A model of inductive bias learning
J. Baxter · 2000
Earlier work this paper cites.
Exploiting task relatedness for multiple task learning
S. Ben-David and R. Schuller · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd, S. P. Boyd, and L. Vandenberghe · 2004
Earlier work this paper cites.
Bounds for linear multi-task learning
A. Maurer · 2006
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Earlier work this paper cites.
Linear thompson sampling revisited
M. Abeille and A. Lazaric · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
Earlier work this paper cites.
Reinforcement learning and optimal control
D. P. Bertsekas · 2019
Earlier work this paper cites.
Sharing knowledge in multi-task deep reinforcement learning
C. D’Eramo, D. Tateo, A. Bonarini, M. Restelli, and J. Peters · 2019
Cited alongside, same era.
Bilinear bandits with low-rank structure
K.-S. Jun, R. Willett, S. Wright, and R. Nowak · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
A. Agarwal, S. Kakade, A. Krishnamurthy, and W. Sun · 2020
Cited alongside, same era.
Few-shot learning via learning the representation, provably
S. S. Du, W. Hu, S. M. Kakade, J. D. Lee, and Q. Lei · 2020
Cited alongside, same era.
Time-uniform chernoff bounds via nonnegative supermartingales
S. R. Howard, A. Ramdas, J. McAuliffe, J. Sekhon, et al · 2020
Cited alongside, same era.
Time-Uniform, Nonparametric, Nonasymptotic Confidence Sequences
S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon · 2021
Later among the works it cites.
Near-optimal representation learning for linear bandits and linear rl
J. Hu, X. Chen, C. Jin, L. Li, and L. Wang · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Later among the works it cites.
Provable representation learning for imitation with contrastive fourier features
O. Nachum and M. Yang · 2021
Later among the works it cites.
Provable meta-learning of linear representations
N. Tripuraneni, C. Jin, and M. I. Jordan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret bound balancing and elimination for model selection in bandits and rl
A. Pacchiano, C. Dann, C. Gentile, and P. Bartlett · 2020
Cited alongside, same era.
On the theory of transfer learning: The importance of task diversity
N. Tripuraneni, M. Jordan, and C. Jin · 2020
Cited alongside, same era.
Impact of representation learning in linear bandits
J. Yang, W. Hu, J. D. Lee, and S. S. Du · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
L. Yang and M. Wang · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
S. S. Du, S. M. Kakade, J. D. Lee, S. Lovett, G. Mahajan, W. Sun, and R. Wang · 2021
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
D. Zhou, Q. Gu, and C. Szepesvari · 2021
Later among the works it cites.
Provable benefits of representational transfer in reinforcement learning
A. Agarwal, Y. Song, W. Sun, K. Wang, M. Wang, and X. Zhang · 2022
Closest in time.
Provable benefit of multitask representation learning in reinforcement learning
Y. Cheng, S. Feng, J. Yang, H. Zhang, and Y. Liang · 2022
Closest in time.
Provable general function class representation learning in multitask bandits and mdps
R. Lu, A. Zhao, S. S. Du, and G. Huang · 2022
Closest in time.
Towards an understanding of default policies in multitask policy optimization
T. Moskovitz, M. Arbel, J. Parker-Holder, and A. Pacchiano · 2022
Closest in time.
Meta learning mdps with linear transition models
R. Müller and A. Pacchiano · 2022
Closest in time.