Fetching the paper…
Reading the bibliography…
We study meta-learning in Markov Decision Processes (MDP) with linear transition models in the undiscounted episodic setting.
Rapid learning or feature reuse? towards understanding the effectiveness of maml
Raghu, A., Raghu, M., Bengio, S., and Vinyals, O. (2019) · 1909
Earlier work this paper cites.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
Lai, T. L., Wei, C. Z., et al. (1982) · 1982
Earlier work this paper cites.
Evolutionary principles in self-referential learning. on learning now to learn: The meta-meta-meta…-hook
Schmidhuber, J. (1987) · 1987
Earlier work this paper cites.
Meta-neural networks that learn by learning
Naik, D. K. and Mammone, R. J. (1992) · 1992
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Learning theory estimates via integral operators and their approximations
Smale, S. and Zhou, D.-X. (2007) · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M. (2008) · 2008
Earlier work this paper cites.
Logarithmic regret for reinforcement learning with linear function approximation
He, J., Zhou, D., and Gu, Q. (2020) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvari, C. (2011) · 2011
Earlier work this paper cites.
Sequential transfer in multi-armed bandit with finite set of models
Azar, M. G., Lazaric, A., and Brunskill, E. (2013) · 2013
Earlier work this paper cites.
Sample complexity of multi-task reinforcement learning
Brunskill, E. and Li, L. (2013) · 2013
Earlier work this paper cites.
Sparse multi-task reinforcement learning
Calandriello, D., Lazaric, A., and Restelli, M. (2015) · 2015
Earlier work this paper cites.
Multi-task learning for contextual bandits
Deshmukh, A. A., Dogan, U., and Scott, C. (2017) · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Learning to learn around a common mean
Denevi, G., Ciliberto, C., Stamos, D., and Pontil, M. (2018) · 2018
Cited alongside, same era.
Learning-to-learn stochastic gradient descent with biased regularization
Denevi, G., Ciliberto, C., Grazzi, R., and Pontil, M. (2019) · 2019
Cited alongside, same era.
Provable guarantees for gradient-based meta-learning
Khodak, M., Balcan, M.-F., and Talwalkar, A. (2019) · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Bandit Algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S. (2020) · 2020
Later among the works it cites.
Taming the herd: Multi-modal meta-learning with a population of agents
Müller, R., Parker-Holder, J., and Pacchiano, A. (2020) · 2020
Later among the works it cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D., He, J., and Gu, Q. (2020) · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S. (2020) · 2020
Later among the works it cites.
A distribution-dependent analysis of meta-learning
Konobeev, M., Kuzborskij, I., and Szepesvári, C. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M. (2019) · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L. F. (2020) · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2020) · 2020
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Cella, L., Lazaric, A., and Pontil, M. (2020) · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Jia, Z., Yang, L., Szepesvari, C., and Wang, M. (2020) · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
On the power of multitask representation learning in linear mdp
Lu, R., Huang, G., and Du, S. S. (2021) · 2021
Later among the works it cites.
When is generalizable reinforcement learning tractable?
Malik, D., Li, Y., and Ravikumar, P. (2021) · 2021
Later among the works it cites.
Provable meta-learning of linear representations
Tripuraneni, N., Jin, C., and Jordan, M. I. (2021) · 2021
Later among the works it cites.
Impact of representation learning in linear bandits
Yang, J., Hu, W., Lee, J. D., and Du, S. S. (2021) · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D., Gu, Q., and Szepesvari, C. (2021) · 2021
Later among the works it cites.