Fetching the paper…
Reading the bibliography…
Can we pre-train a generalist agent from a large amount of unlabeled offline trajectories such that it can be immediately adapted to any new downstream tasks in a zero-shot manner? In this work, we present a functional reward encoding (FRE) as a general, scalable solution to this zero-shot RL problem.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 2000
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Li, L., Yang, R., and Luo, D · 2010
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A., Di Castro, D., and Mannor, S · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Earlier work this paper cites.
Learning multi-level hierarchies with hindsight
Levy, A., Konidaris, G., Platt, R., and Saenko, K · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Universal successor features approximators
Borsa, D., Barreto, A., Quan, J., Mankowitz, D., Munos, R., Van Hasselt, H., Silver, D., and Schaul, T · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Dher: Hindsight experience replay for dynamic goals
Fang, M., Zhou, C., Shi, B., Gong, B., Xu, J., and Zhang, T · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Cited alongside, same era.
Semi-parametric topological memory for navigation
Savinov, N., Dosovitskiy, A., and Koltun, V · 2018
Cited alongside, same era.
Lancon-learn: Learning with language to enable generalization in multi-task manipulation
Silva, A., Moorman, N., Silva, W., Zaidi, Z., Gopalan, N., and Gombolay, M · 2021
Later among the works it cites.
Multi-task reinforcement learning with context-based representations
Sodhani, S., Zhang, A., and Pineau, J · 2021
Later among the works it cites.
Learning more skills through optimistic exploration
Strouse, D., Baumli, K., Warde-Farley, D., Mnih, V., and Hansen, S · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
Touati, A. and Ollivier, Y · 2021
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
Eysenbach, B., Zhang, T., Levine, S., and Salakhutdinov, R. R · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Kim, H., Mnih, A., Schwarz, J., Garnelo, M., Eslami, A., Rosenbaum, D., Vinyals, O., and Teh, Y. W · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Cited alongside, same era.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P · 2022
Later among the works it cites.
Hierarchical planning through goal-conditioned offline reinforcement learning
Li, J., Tang, C., Tomizuka, M., and Zhan, W · 2022
Later among the works it cites.
Offline meta-reinforcement learning with online self-supervision
Pong, V. H., Nair, A. V., Smith, L. M., Huang, C., and Levine, S · 2022
Later among the works it cites.
Does zero-shot reinforcement learning exist?
Touati, A., Rapin, J., and Ollivier, Y · 2022
Later among the works it cites.
Rethinking goal-conditioned supervised learning and its connection to offline rl
Yang, R., Lu, Y., Li, W., Sun, H., Fang, M., Du, Y., Li, X., Han, L., and Zhang, C · 2022
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
Yarats, D., Brandfonbrener, D., Liu, H., Laskin, M., Abbeel, P., Lazaric, A., and Pinto, L · 2022
Later among the works it cites.
Robust task representations for offline meta-reinforcement learning via contrastive learning
Yuan, H. and Lu, Z · 2022
Later among the works it cites.
f f -policy gradients: A general framework for goal conditioned rl using f f -divergences
Agarwal, S., Durugkar, I., Stone, P., and Zhang, A · 2023
Later among the works it cites.
Self-supervised reinforcement learning that transfers using random features
Chen, B., Zhu, C., Agrawal, P., Zhang, K., and Gupta, A · 2023
Later among the works it cites.
Unsupervised behavior extraction via random intent priors
Hu, H., Yang, Y., Ye, J., Mai, Z., and Zhang, C · 2023
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Optimal goal-reaching reinforcement learning via quasimetric learning
Wang, T., Torralba, A., Isola, P., and Zhang, A · 2023
Later among the works it cites.