Fetching the paper…
Reading the bibliography…
Meta reinforcement learning (meta-RL) extracts knowledge from previous tasks and achieves fast adaptation to new tasks.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
On learning how to learn learning strategies
Schmidhuber, J · 1995
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Storck, J., Hochreiter, S., and Schmidhuber, J · 1995
Earlier work this paper cites.
What’s interesting?
Schmidhuber, J · 1997
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Sun, Y., Gomez, F., and Schmidhuber, J · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A., and Bengio, Y · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Cited alongside, same era.
Thompson sampling is asymptotically optimal in general environments
Leike, J., Lattimore, T., Orseau, L., and Hutter, M · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Meta-learning: from few-shot learning to rapid reinforcement learning
Finn, C. and Levine, S · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N · 2019
Later among the works it cites.
Emi: Exploration with mutual information
Kim, H., Kim, J., Jeong, Y., Levine, S., and Song, H. O · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Cited alongside, same era.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Environment probing interaction policies
Zhou, W., Pinto, L., and Gupta, A · 2018
Cited alongside, same era.
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Later among the works it cites.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., Abbeel, P., Shingarey, D., Kaul, L., Asfour, T., Athanasios, C. D., Zhou, Y., et al · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2019
Later among the works it cites.
Mame: Model-agnostic meta-exploration
Gurumurthy, S., Kumar, S., and Sycara, K · 2020
Closest in time.
Context-aware dynamics model for generalization in model-based reinforcement learning
Lee, K., Seo, Y., Lee, S., Lee, H., and Shin, J · 2020
Closest in time.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Closest in time.