Fetching the paper…
Reading the bibliography…
In meta reinforcement learning (meta RL), an agent seeks a Bayes-optimal policy -- the optimal policy when facing an unknown task that is sampled from some known task distribution.
Multi-task batch reinforcement learning with metric learning
Li, J., Vuong, Q., Liu, S., Liu, M., Ciosek, K., Christensen, H. I., and Su, H · 1909
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 1995
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Duff, M. O · 2002
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2004
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Li, L., Yang, R., and Luo, D · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
RL 2 \text{RL}^{2} : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T · 2018
Cited alongside, same era.
Neural predictive belief representations
Guo, Z. D., Azar, M. G., Piot, B., Pires, B. A., and Munos, R · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Later among the works it cites.
Approximate information state for approximate planning and reinforcement learning in partially observed systems
Subramanian, J., Sinha, A., Seraj, R., and Mahajan, A · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2020
Later among the works it cites.
Offline meta reinforcement learning–identifiability challenges and effective data collection strategies
Dorfman, R., Shenfeld, I., and Tamar, A · 2021
Later among the works it cites.
panda-gym: Open-Source Goal-Conditioned Environments for Robotic Learning
Gallouédec, Q., Cazin, N., Dellandréa, E., and Chen, L · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Fakoor, R., Chaudhari, P., Soatto, S., and Smola, A. J · 2019
Cited alongside, same era.
Meta reinforcement learning as task inference
Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N · 2019
Cited alongside, same era.
Meta-learning of sequential strategies
Ortega, P. A., Wang, J. X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., et al · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Cited alongside, same era.
Han, T., Huang, H., Yang, Z., and Han, W · 2021
Later among the works it cites.
Provably improved context-based offline meta-rl with attention and contrastive learning
Li, L., Huang, Y., Chen, M., Luo, S., Luo, D., and Huang, J · 2021
Later among the works it cites.
Unsupervised feature learning for manipulation with contrastive domain randomization
Rabinovitz, C., Grupen, N., and Tamar, A · 2021
Later among the works it cites.
Automatic data augmentation for generalization in reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R · 2021
Later among the works it cites.
Wang, B., Xu, S., Keutzer, K., Gao, Y., and Wu, B · 2021
Later among the works it cites.
Exploration in approximate hyper-state space for meta reinforcement learning
Zintgraf, L. M., Feng, L., Lu, C., Igl, M., Hartikainen, K., Hofmann, K., and Whiteson, S · 2021
Later among the works it cites.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P · 2022
Later among the works it cites.
Recurrent model-free rl can be a strong baseline for many pomdps
Ni, T., Eysenbach, B., and Salakhutdinov, R · 2022
Later among the works it cites.