Fetching the paper…
Reading the bibliography…
Humans can leverage prior experience and learn novel tasks from a handful of demonstrations.
Towards one shot learning by imitation for humanoid robots
Wu, Y. and Demiris, Y · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Learning feed-forward one-shot learners
Bertinetto, L., Henriques, J. F., Valmadre, J., Torr, P., and Vedaldi, A · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Earlier work this paper cites.
Low data drug discovery with one-shot learning
Altae-Tran, H., Ramsundar, B., Pappu, A. S., and Pande, V · 2017
Earlier work this paper cites.
Smash: one-shot model architecture search through hypernetworks
Brock, A., Lim, T., Ritchie, J. M., and Weston, N · 2017
Earlier work this paper cites.
Duan, Y., Andrychowicz, M., Stadie, B. C., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Fast parameter adaptation for few-shot image captioning and visual question answering
Dong, X., Zhu, L., Zhang, D., Yang, Y., and Wu, F · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
On First-Order Meta-Learning Algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Taco: Learning task decomposition via temporal alignment for control
Shiarlis, K., Wulfmeier, M., Salter, S., Whiteson, S., and Posner, I · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D · 2020
Later among the works it cites.
Generalizing from a few examples: A survey on few-shot learning
Wang, Y., Yao, Q., Kwok, J. T., and Ni, L. M · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 2018
Cited alongside, same era.
Off-Policy Policy Gradient Algorithms by Constraining the State Distribution Shift
Islam, R., Teru, K. K., Sharma, D., and Pineau, J · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Meta-Learning with Implicit Gradients
Rajeswaran, A., Finn, C., Kakade, S. M., and Levine, S · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S · 2021
Later among the works it cites.
Reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Badgr: An autonomous self-supervised learning-based navigation system
Kahn, G., Abbeel, P., and Levine, S · 2021
Later among the works it cites.
MOReL : Model-Based Offline Reinforcement Learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2021
Later among the works it cites.
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2021
Later among the works it cites.
Offline meta-reinforcement learning with advantage weighting
Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C · 2021
Later among the works it cites.
COMBO: Conservative Offline Model-Based Policy Optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.