Fetching the paper…
Reading the bibliography…
In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL).
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2017
Earlier work this paper cites.
Embodied question answering
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D · 2018
Earlier work this paper cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
Frostig, R., Johnson, M. J., and Leary, C · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Earlier work this paper cites.
Learning from synthetic data: Addressing domain shift for semantic segmentation
Sankaranarayanan, S., Balaji, Y., Jain, A., Lim, S. N., and Chellappa, R · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Solving the rubik’s cube with deep reinforcement learning and search
Agostinelli, F., McAleer, S., Shmakov, A., and Baldi, P · 2019
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J. W., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Earlier work this paper cites.
Provably efficient RL with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Earlier work this paper cites.
Defense against adversarial attacks using feature scattering-based adversarial training
Zhang, H. and Wang, J · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2020
Earlier work this paper cites.
Active domain randomization
Mehta, B., Diaz, M., Golemo, F., Pal, C. J., and Paull, L · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
Zhang, K., Kakade, S., Basar, T., and Yang, L · 2020
Cited alongside, same era.
AMP: adversarial motion priors for stylized physics-based character control
Peng, X. B., Ma, Z., Abbeel, P., Levine, S., and Kanazawa, A · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Vision-language models provide promptable representations for reinforcement learning
Chen, W., Mees, O., Kumar, A., and Levine, S · 2023
Later among the works it cites.
CLIP-Motion: Learning reward functions for robotic actions using consecutive observations
Dang, X., Edelkamp, S., and Ribault, N · 2023
Later among the works it cites.
Vision-language models as success detectors
Du, Y., Konyushkova, K., Denil, M., Raju, A., Landon, J., Hill, F., de Freitas, N., and Cabi, S · 2023
Later among the works it cites.
Physically grounded vision-language models for robotic manipulation
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2023
Later among the works it cites.
Toward general-purpose robots via foundation models: A survey and meta-analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S · 2021
Cited alongside, same era.
Multi-task reinforcement learning with context-based representations
Sodhani, S., Zhang, A., and Pineau, J · 2021
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2021
Cited alongside, same era.
Adaptive risk minimization: Learning to adapt to domain shift
Zhang, M., Marklund, H., Dhawan, N., Gupta, A., Levine, S., and Finn, C · 2021
Cited alongside, same era.
Do as I can, not as I say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Bińkowski, M. a., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Hu, Y., Xie, Q., Jain, V., Francis, J., Patrikar, J., Keetha, N., Kim, S., Xie, Y., Zhang, T., Zhao, S., Chong, Y. Q., Wang, C., Sycara, K., Johnson-Roberson, M., Batra, D., Wang, X., Scherer, S., Kira, Z., Xia, F., and Bisk, Y · 2023
Later among the works it cites.
Grounded decoding: Guiding text generation with grounded models for embodied agents
Huang, W., Xia, F., Shah, D., Driess, D., Zeng, A., Lu, Y., Florence, P., Mordatch, I., Levine, S., Hausman, K., and brian ichter · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M · 2023
Later among the works it cites.
Can agents run relay race with strangers? generalization of RL to out-of-distribution trajectories
Lan, L.-C., Zhang, H., and Hsieh, C.-J · 2023
Later among the works it cites.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Later among the works it cites.
FoMo rewards: Can we cast foundation models as reward functions?
Lubana, E. S., Brehmer, J., de Haan, P., and Cohen, T · 2023
Later among the works it cites.
R3M: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2023
Later among the works it cites.
Lift: Unsupervised reinforcement learning with foundation models as teachers
Nam, T., Lee, J., Zhang, J., Hwang, S. J., Lim, J. J., and Pertsch, K · 2023
Later among the works it cites.
Roboclip: One demonstration is enough to learn robot policies
Sontakke, S. A., Arnold, S., Zhang, J., Pertsch, K., Biyik, E., Sadigh, D., Finn, C., and Itti, L · 2023
Later among the works it cites.
Distilling internet-scale vision-language models into embodied agents
Sumers, T., Marino, K., Ahuja, A., Fergus, R., and Dasgupta, I · 2023
Later among the works it cites.
Jump-start reinforcement learning
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., et al · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
LLM lies: Hallucinations are not bugs, but features as adversarial examples
Yao, J.-Y., Ning, K.-P., Liu, Z.-H., Ning, M.-N., and Yuan, L · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Arenas, M. G., Chiang, H.-T. L., Erez, T., Hasenclever, L., Humplik, J., Ichter, B., Xiao, T., Xu, P., Zeng, A., Zhang, T., Heess, N., Sadigh, D., Tan, J., Tassa, Y., and Xia, F · 2023
Later among the works it cites.
Chakraborty, N., Ornik, M., and Driggs-Campbell, K · 2024
Closest in time.