Fetching the paper…
Reading the bibliography…
We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning.
Design of experiments
R. A. Fisher · 1936
Earlier work this paper cites.
Design of a low cost, general purpose robot
M. H. Smith and L. S. Coles · 1973
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
More than a million ways to be pushed. a high-fidelity experimental dataset of planar pushing
K.-T. Yu, M. Bauza, N. Fazeli, and A. Rodriguez · 2016
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
D. Cer, Y. Yang, S. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, Y. Sung, B. Strope, and R. Kurzweil · 2018
Earlier work this paper cites.
Task-embedded control networks for few-shot imitation learning
S. James, M. Bloesch, and A. J. Davison · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, A. Xie, S. Dasari, T. Zhang, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Earlier work this paper cites.
Skew-fit: State-covering self-supervised reinforcement learning
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Self-supervised policy adaptation during deployment
N. Hansen, R. Jangir, Y. Sun, G. Alenyà, P. Abbeel, A. A. Efros, L. Pinto, and X. Wang · 2020
Earlier work this paper cites.
Human instruction-following with deep reinforcement learning via transfer-learning from text
F. Hill, S. Mokra, N. Wong, and T. Harley · 2020
Earlier work this paper cites.
The foundation of efficient robot learning
L. P. Kaelbling · 2020
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Transformers for one-shot visual imitation
S. Dasari and A. Gupta · 2021
Earlier work this paper cites.
Open-vocabulary image segmentation
G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin · 2021
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2021
Cited alongside, same era.
The surprising effectiveness of representation learning for visual imitation
J. Pari, N. M. Shafiullah, S. P. Arunachalam, and L. Pinto · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Tokenlearner: Adaptive space-time tokenization for videos
M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova · 2021
Cited alongside, same era.
Formal mathematics statement curriculum learning
S. Polu, J. M. Han, K. Zheng, M. Baksys, I. Babuschkin, and I. Sutskever · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, and L. Wang · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
Scaling vision transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rrl: Resnet as representation for reinforcement learning
R. Shah and V. Kumar · 2021
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Cited alongside, same era.
Visual imitation made easy
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto · 2021
Cited alongside, same era.
Do as I can, not as I say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
From play to policy: Conditional behavior generation from uncurated robot data
Z. J. Cui, Y. Wang, N. Muhammad, L. Pinto, et al · 2022
Cited alongside, same era.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters, 2023
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. van Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. P. Collier, A. Gritsenko, V. Birodkar, C. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetić, D. Tran, T. Kipf, M. Lučić, X. Zhai, D. Keysers, J. Harmsen, and N. Houlsby · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
Language is not all you need: Aligning perception with language models
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu, et al · 2023
Closest in time.
Language-driven representation learning for robotics
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang · 2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Liv: Language-image representations and rewards for robotic control
Y. J. Ma, W. Liang, V. Som, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P. Luo · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, b. ichter, and S. Levine · 2023
Closest in time.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, et al · 2023
Closest in time.
Distilling internet-scale vision-language models into embodied agents
T. Sumers, K. Marino, A. Ahuja, R. Fergus, and I. Dasgupta · 2023
Closest in time.
Ul2: Unifying language learning paradigms, 2023
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, S. Shakeri, D. Bahri, T. Schuster, H. S. Zheng, D. Zhou, N. Houlsby, and D. Metzler · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor · 2023
Closest in time.
Symbol tuning improves in-context learning in language models, 2023
J. Wei, L. Hou, A. Lampinen, X. Chen, D. Huang, Y. Tay, X. Chen, Y. Lu, D. Zhou, T. Ma, and Q. V. Le · 2023
Closest in time.
Tidybot: Personalized robot assistance with large language models
J. Wu, R. Antonova, A. Kan, M. Lepert, A. Zeng, S. Song, J. Bohg, S. Rusinkiewicz, and T. Funkhouser · 2023
Closest in time.
Grounding classical task planners via vision-language models
X. Zhang, Y. Ding, S. Amiri, H. Yang, A. Kaminski, C. Esselink, and S. Zhang · 2023
Closest in time.