Fetching the paper…
Reading the bibliography…
Consider the following instance of the Offline Meta Reinforcement Learning (OMRL) problem: given the complete training logs of $N$ conventional RL agents, trained on $N$ different tasks, design a meta-agent that can quickly maximize reward in a new, unseen task from the same task distribution.
Acting optimally in partially observable stochastic domains
Cassandra, A. R., Kaelbling, L. P., and Littman, M. L · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 1995
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Duff, M. O · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
RL 2 \text{RL}^{2} : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Earlier work this paper cites.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2018
Cited alongside, same era.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T · 2018
Cited alongside, same era.
Neural predictive belief representations
Guo, Z. D., Azar, M. G., Piot, B., Pires, B. A., and Munos, R · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Meta reinforcement learning as task inference
Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N · 2019
Later among the works it cites.
Meta-learning of sequential strategies
Ortega, P. A., Wang, J. X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., et al · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Later among the works it cites.
Simultaneously learning vision and feature-based control policies for real-world ball-in-a-cup
Schwab, D., Springenberg, T., Martins, M. F., Lampe, T., Neunert, M., Abdolmaleki, A., Hertweck, T., Hafner, R., Nori, F., and Riedmiller, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Bayesian policy optimization for model uncertainty
Lee, G., Hou, B., Mandalika, A., Lee, J., Choudhury, S., and Srinivasa, S. S · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Safe policy learning from observations
Sarafian, E., Tamar, A., and Kraus, S · 2018
Cited alongside, same era.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Cited alongside, same era.
Fakoor, R., Chaudhari, P., Soatto, S., and Smola, A. J · 2019
Cited alongside, same era.
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Closest in time.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Closest in time.
Multi-task batch reinforcement learning with metric learning
Li, J., Vuong, Q., Liu, S., Liu, M., Ciosek, K., Christensen, H. I., and Su, H · 2020
Closest in time.
Offline meta-reinforcement learning with advantage weighting
Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C · 2020
Closest in time.
Meta-reinforcement learning for robotic industrial insertion tasks
Schoettler, G., Nair, A., Ojea, J. A., Levine, S., and Solowjow, E · 2020
Closest in time.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2020
Closest in time.