Fetching the paper…
Reading the bibliography…
Recent advances in generalist robot manipulation leverage pre-trained Vision-Language Models (VLMs) and large-scale robot demonstrations to tackle diverse tasks in a zero-shot manner.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, J. Ma, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video, 2022
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Radosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, D. Batra, V. Cartillier, S. Crane, T. Do, M. Doulaty, A. Erapalli, C. Feichtenhofer, A. Fragomeni, Q. Fu, A. Gebreselasie, C. Gonzalez, J. Hillis, X. Huang, Y. Huang, W. Jia, W. Khoo, J. Kolar, S. Kottur, A. Kumar, F. Landini, C. Li, Y. Li, Z. Li, K. Mangalam, R. Modhugu, J. Munro, T. Murrell, T. Nishiyasu, W. Price, P. R. Puentes, M. Ramazanova, L. Sari, K. Somasundaram, A. Southerland, Y. Sugano, R. Tao, M. Vo, Y. Wang, X. Wu, T. Yagi, Z. Zhao, Y. Zhu, P. Arbelaez, D. Crandall, D. Damen, G. M. Farinella, C. Fuegen, B. Ghanem, V. K. Ithapu, C. V. Jawahar, H. Joo, K. Kitani, H. Li, R. Newcombe, A. Oliva, H. S. Park, J. M. Rehg, Y. Sato, J. Shi, M. Z. Shou, A. Torralba, L. Torresani, M. Yan, and J. Malik · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2023
Earlier work this paper cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Earlier work this paper cites.
VIP: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2023
Earlier work this paper cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, T. Wu, J. Vakil, P. Abbeel, J. Malik, D. Batra, Y. Lin, O. Maksymets, A. Rajeswaran, and F. Meier · 2023
Earlier work this paper cites.
Learning universal policies via text-guided video generation, 2023
Y. Du, M. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Earlier work this paper cites.
Zero-shot robot manipulation from passive human videos, 2023
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Earlier work this paper cites.
Pali-3 vision language models: Smaller, faster, stronger, 2023
X. Chen, X. Wang, L. Beyer, A. Kolesnikov, J. Wu, P. Voigtlaender, B. Mustafa, S. Goodman, I. Alabdulmohsin, P. Padlewski, D. Salz, X. Xiong, D. Vlasic, F. Pavetic, K. Rong, T. Yu, D. Keysers, X. Zhai, and R. Soricut · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Earlier work this paper cites.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Cited alongside, same era.
Flow matching for generative modeling
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2023
Cited alongside, same era.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot, 2023
H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu · 2023
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, A. Lee, K. Fang, C. Finn, and S. Levine · 2023
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos, 2024
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Later among the works it cites.
Bootstap: Bootstrapped training for tracking-any-point
C. Doersch, P. Luc, Y. Yang, D. Gokay, S. Koppula, A. Gupta, J. Heyward, I. Rocco, R. Goroshin, J. a. Carreira, and A. Zisserman · 2024
Later among the works it cites.
R. Wang, S. Xu, C. Dai, J. Xiang, Y. Deng, X. Tong, and J. Yang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y. Deng, S. Xu, Y. Zhang, X. Wang, B. Liu, J. Fu, J. Bao, D. Chen, Y. Shi, J. Yang, and B. Guo · 2024
Cited alongside, same era.
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation, 2024
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani · 2024
Cited alongside, same era.
Flow as the cross-domain manipulation interface
M. Xu, Z. Xu, Y. Xu, C. Chi, G. Wetzstein, M. Veloso, and S. Song · 2024
Cited alongside, same era.
Any-point trajectory modeling for policy learning, 2024
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2024
Cited alongside, same era.
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani · 2024
Cited alongside, same era.
Egomimic: Scaling imitation learning via egocentric video, 2024
S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu · 2024
Cited alongside, same era.
Hand-object interaction pretraining from videos
H. G. Singh, A. Loquercio, C. Sferrazza, J. Wu, H. Qi, P. Abbeel, and J. Malik · 2024
Cited alongside, same era.
W. Liang, L. Yu, L. Luo, S. Iyer, N. Dong, C. Zhou, G. Ghosh, M. Lewis, W. tau Yih, L. Zettlemoyer, and X. V. Lin · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Shahriari, C. L. Lan, C. A. Choquette-Choo, C. Crepy, D. Cer, D. Ippolito, D. Reid, E. Buchatskaya, E. Ni, E. Noland, G. Yan, G. Tucker, G.-C. Muraru, G. Rozhdestvenskiy, H. Michalewski, I. Tenney, I. Grishchenko, J. Austin, J. Keeling, J. Labanowski, J.-B. Lespiau, J. Stanway, J. Brennan, J. Chen, J. Ferret, J. Chiu, and et al · 2024
Later among the works it cites.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao · 2024
Later among the works it cites.
Monst3r: A simple approach for estimating geometry in the presence of motion, 2024
J. Zhang, C. Herrmann, J. Hur, V. Jampani, T. Darrell, F. Cole, D. Sun, and M.-H. Yang · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y. Liu, D. Xiang, G. Wetzstein, and T.-Y. Lin · 2025
Closest in time.
Revla: Reverting visual domain limitation of robotic foundation models, 2025
S. Dey, J.-N. Zaech, N. Nikolov, L. V. Gool, and D. P. Paudel · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model, 2025
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, and X. Li · 2025
Closest in time.
Vilp: Imitation learning with latent video planning
Z. Xu, Q. Qiu, and Y. She · 2025
Closest in time.
Point policy: Unifying observations and actions with key points for robot manipulation, 2025
S. Haldar and L. Pinto · 2025
Closest in time.
J. Ren, P. Sundaresan, D. Sadigh, S. Choudhury, and J. Bohg · 2025
Closest in time.
Pre-training auto-regressive robotic models with 4d representations, 2025
D. Niu, Y. Sharma, H. Xue, G. Biamby, J. Zhang, Z. Ji, T. Darrell, and R. Herzig · 2025
Closest in time.
Latent action pretraining from videos
S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo · 2025
Closest in time.