Fetching the paper…
Reading the bibliography…
Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding.
A tutorial on the cross-entropy method
De Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y · 2005
Earlier work this paper cites.
Open source computer vision library
Itseez · 2015
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
Zeng, A., Florence, P., Tompson, J., Welker, S., Chien, J., Attarian, M., Armstrong, T., Krasin, I., Duong, D., Sindhwani, V., et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Can foundation models perform zero-shot task specification for robot manipulation?
Cui, Y., Niekum, S., Gupta, A., Kumar, V., and Rajeswaran, A · 2022
Earlier work this paper cites.
Clip-nav: Using clip for zero-shot vision-and-language navigation
Dorbala, V. S., Sigurdsson, G., Piramuthu, R., Thomason, J., and Sukhatme, G. S · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Planning with large language models via corrective re-prompting
Raman, S. S., Cohen, V., Rosen, E., Idrees, I., Paulius, D., and Tellex, S · 2022
Earlier work this paper cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Gps: Genetic prompt search for efficient few-shot learning
Xu, H., Chen, Y., Du, Y., Shao, N., Wang, Y., Li, H., and Yang, Z · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., et al · 2022
Cited alongside, same era.
Large language models are human-level prompt engineers
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J · 2022
Text2motion: From natural language instructions to feasible plans
Lin, K., Agia, C., Migimatsu, T., Pavone, M., and Bohg, J · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Large language models as general pattern machines
Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Making large multimodal models understand arbitrary visual prompts
Cai, M., Liu, H., Mustikovela, S. K., Meyer, G. P., Chai, Y., Park, D., and Lee, Y. J · 2023
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning
Chen, B., Xia, F., Ichter, B., Rao, K., Gopalakrishnan, K., Ryoo, M. S., Stone, A., and Kappler, D · 2023
Cited alongside, same era.
Foundation models in robotics: Applications, challenges, and the future
Firoozi, R., Tucker, J., Tian, S., Majumdar, A., Sun, J., Liu, W., Zhu, Y., Song, S., Kapoor, A., Hausman, K., et al · 2023
Cited alongside, same era.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
Gadre, S. Y., Wortsman, M., Ilharco, G., Schmidt, L., and Song, S · 2023
Cited alongside, same era.
Physically grounded vision-language models for robotic manipulation
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini, T., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Automatic prompt optimization with" gradient descent" and beam search
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M · 2023
Later among the works it cites.
What does clip know about a red circle? visual prompt engineering for vlms
Shtedritski, A., Rupprecht, C., and Vedaldi, A · 2023
Later among the works it cites.
Generalized planning in pddl domains with pretrained large language models
Silver, T., Dan, S., Srinivas, K., Tenenbaum, J. B., Kaelbling, L. P., and Katz, M · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Later among the works it cites.
On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving
Wen, L., Yang, X., Fu, D., Wang, X., Cai, P., Li, X., Ma, T., Li, Y., Xu, L., Shang, D., et al · 2023
Later among the works it cites.
Tidybot: Personalized robot assistance with large language models
Wu, J., Antonova, R., Kan, A., Lepert, M., Zeng, A., Song, S., Bohg, J., Rusinkiewicz, S., and Funkhouser, T · 2023
Later among the works it cites.
Xu, J., Zhou, X., Yan, S., Gu, X., Arnab, A., Sun, C., Wang, X., and Schmid, C · 2023
Later among the works it cites.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
Yan, A., Yang, Z., Zhu, W., Lin, K., Li, L., Wang, J., Yang, J., Zhong, Y., McAuley, J., Gao, J., et al · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Arenas, M. G., Chiang, H.-T. L., Erez, T., Hasenclever, L., Humplik, J., et al · 2023
Later among the works it cites.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Chen, B., Xu, Z., Kirmani, S., Ichter, B., Driess, D., Florence, P., Sadigh, D., Guibas, L., and Xia, F · 2024
Closest in time.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.