Fetching the paper…
Reading the bibliography…
Meta-reinforcement learning (RL) methods can meta-train policies that adapt to new tasks with orders of magnitude less data than standard RL, but meta-training itself is costly and time-consuming.
Benchmarking batch deep reinforcement learning algorithms
Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J · 1910
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Teleoperation of a robot manipulator using a vision-based human-robot interface
Kofman, J., Wu, X., Luu, T. J., and Verma, S · 2005
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Universal Value Function Approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
End-to-end goal-driven web navigation
Nogueira, R. and Cho, K · 2016
Earlier work this paper cites.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., Mcgrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Earlier work this paper cites.
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A · 2017
Earlier work this paper cites.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Levine, S., Pastor, P., Krizhevsky, A., Ibarz, J., and Quillen, D · 2018
Cited alongside, same era.
Visual Reinforcement Learning with Imagined Goals
Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Cited alongside, same era.
Replab: A reproducible low-cost arm benchmark platform for robotic learning
Yang, B., Zhang, J., Pong, V., Levine, S., and Jayaraman, D · 2019
Later among the works it cites.
Offline meta reinforcement learning
Dorfman, R. and Tamar, A · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Kamienny, P.-A., Pirotta, M., Lazaric, A., Lavril, T., Usunier, N., and Denoyer, L · 2020
Later among the works it cites.
Semi-supervised reward learning for offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration
Péré, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y · 2018
Cited alongside, same era.
Temporal Difference Models: Model-Free Deep RL For Model-Based Control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Deep reinforcement learning for general video game ai
Torrado, R. R., Bontrager, P., Togelius, J., Liu, J., and Perez-Liebana, D · 2018
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D · 2018
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Barreto, A., Borsa, D., Quan, J., Schaul, T., Silver, D., Hessel, M., Mankowitz, D., Žídek, A., and Munos, R · 2019
Cited alongside, same era.
Konyushkova, K., Zolna, K., Aytar, Y., Novikov, A., Reed, S., Cabi, S., and de Freitas, N · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Later among the works it cites.
Learning agile robotic locomotion skills by imitating animals
Peng, X. B., Coumans, E., Zhang, T., Lee, T.-W., Tan, J., and Levine, S · 2020
Later among the works it cites.
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H., Hessel, M., Oh, J., Singh, S., and Silver, D · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Later among the works it cites.
Meld: Meta-reinforcement learning from images via latent state models
Zhao, T. Z., Nagabandi, A., Rakelly, K., Finn, C., and Levine, S · 2020
Later among the works it cites.
Varibad: a very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
Zolna, K., Novikov, A., Konyushkova, K., Gulcehre, C., Wang, Z., Aytar, Y., Denil, M., de Freitas, N., and Reed, S · 2020
Later among the works it cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Coumans, E. and Bai, Y · 2021
Closest in time.
What can i do here? learning new skills by imagining visual affordances
Khazatsky, A., Nair, A., Jing, D., and Levine, S · 2021
Closest in time.
Offline meta-reinforcement learning with advantage weighting
Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C · 2021
Closest in time.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Closest in time.
Towards real robot learning in the wild: A case study in bipedal locomotion
Bloesch, M., Humplik, J., Patraucean, V., Hafner, R., Haarnoja, T., Byravan, A., Siegel, N. Y., Tunyasuvunakool, S., Casarini, F., Batchelor, N., et al · 2022
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2062
Closest in time.