Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train.
Recognizing indoor scenes
A. Quattoni and A. Torralba · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2016
Earlier work this paper cites.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. Stadie, O. Jonathan Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Learning to navigate in cities without a map
P. W. Mirowski, M. K. Grimes, M. Malinowski, K. M. Hermann, K. Anderson, D. Teplyashin, K. Simonyan, K. Kavukcuoglu, A. Zisserman, and R. Hadsell · 2018
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei · 2018
Earlier work this paper cites.
Task-embedded control networks for few-shot imitation learning
S. James, M. Bloesch, and A. J. Davison · 2018
Earlier work this paper cites.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra · 2019
Earlier work this paper cites.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Earlier work this paper cites.
Rearrangement: A challenge for embodied ai
D. Batra, A. X. Chang, S. Chernova, A. J. Davison, J. Deng, V. Koltun, S. Levine, J. Malik, I. Mordatch, R. Mottaghi, M. Savva, and H. Su · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. Rovick Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. A. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Kubric: A scalable dataset generator
K. Greff, F. Belletti, L. Beyer, C. Doersch, Y. Du, D. Duckworth, D. J. Fleet, D. Gnanapragasam, F. Golemo, C. Herrmann, T. Kipf, A. Kundu, D. Lagun, I. H. Laradji, H.-T. Liu, H. Meyer, Y. Miao, D. Nowrouzezahrai, C. Oztireli, E. Pot, N. Radwan, D. Rebain, S. Sabour, M. S. M. Sajjadi, M. Sela, V. Sitzmann, A. Stone, D. Sun, S. Vora, Z. Wang, T. Wu, K. M. Yi, F. Zhong, and A. Tagliasacchi · 2022
Earlier work this paper cites.
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
H. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu · 2022
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. H. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence · 2023
Earlier work this paper cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Earlier work this paper cites.
Criteria: a new benchmarking paradigm for evaluating trajectory prediction models for autonomous driving
C. Chen, M. Pourkeshavarz, and A. Rasouli · 2023
Cited alongside, same era.
Behavior retrieval: Few-shot imitation learning by querying unlabeled datasets
M. Du, S. Nair, D. Sadigh, and C. Finn · 2023
Cited alongside, same era.
All in tokens: Unifying output space of visual tasks via soft token
J. Ning, C. Li, Z. Zhang, Z. Geng, Q. Dai, K. He, and H. Hu · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi · 2023
Cited alongside, same era.
Octo: An open-source generalist robot policy
D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Cited alongside, same era.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. H. Vuong, and T. Xiao · 2024
Later among the works it cites.
Z. Wang, Z. Zhou, J. Song, Y. Huang, Z. Shu, and L. Ma · 2024
Later among the works it cites.
Covla: Comprehensive vision-language-action dataset for autonomous driving
H. Arai, K. Miwa, K. Sasaki, Y. Yamaguchi, K. Watanabe, S. Aoki, and I. Yamamoto · 2024
Later among the works it cites.
Flowretrieval: Flow-guided data retrieval for few-shot imitation learning
L.-H. Lin, Y. Cui, A. Xie, T. Hua, and D. Sadigh · 2024
Later among the works it cites.
One-shot imitation learning with invariance matching for robotic manipulation
X. Zhang and A. Boularias · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al · 2024
Cited alongside, same era.
Towards generalist robot policies: What matters in building vision-language-action models
X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu · 2024
Cited alongside, same era.
Maniskill3: Gpu parallelized robotics simulation and rendering for generalizable embodied ai
S. Tao, F. Xiang, A. Shukla, Y. Qin, X. Hinrichsen, X. Yuan, C. Bao, X. Lin, Y. Liu, T. kai Chan, Y. Gao, X. Li, T. Mu, N. Xiao, A. Gurha, Z. Huang, R. Calandra, R. Chen, S. Luo, and H. Su · 2024
Cited alongside, same era.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
R. Zheng, Y. Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang · 2024
Cited alongside, same era.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y. Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y. Chou, A. Chheda, J. Sparks, S. Skjonsberg, M. Schmitz, A. Sarnat, B. Bischoff, P. Walsh, C. Newell, P. Wolters, T. Gupta, K.-H. Zeng, J. Borchardt, D. Groeneveld, J. Dumas, C. Nam, S. Lebrecht, C. M. Wittlif, C. Schoenick, O. Michel, R. Krishna, L. Weihs, N. A. Smith, H. Hajishirzi, R. Girshick, A. Farhadi, and A. Kembhavi · 2024
Cited alongside, same era.
Later among the works it cites.
Ditto: Demonstration imitation by trajectory transformation
N. Heppert, M. Argus, T. Welschehold, T. Brox, and A. Valada · 2024
Later among the works it cites.
Keypoint action tokens enable in-context imitation learning in robotics
N. Di Palo and E. Johns · 2024
Later among the works it cites.
In-context imitation learning via next-token prediction
L. Fu, H. Huang, G. Datta, L. Y. Chen, W. C.-H. Panitch, F. Liu, H. Li, and K. Goldberg · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, S. Qin, R. R. Ingle, E. Bugliarello, S. Kazemzadeh, T. Mesnard, I. M. Alabdulmohsin, L. Beyer, and X.-Q. Zhai · 2024
Later among the works it cites.
Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y. Kim, W. Han, A. Herrasti, et al · 2024
Later among the works it cites.
Variational distillation of diffusion policies into mixture of experts
H. Zhou, D. Blessing, G. Li, O. Celik, X. Jia, G. Neumann, and R. Lioutikov · 2024
Later among the works it cites.
A0: An affordance-aware hierarchical model for general robotic manipulation
R. Xu, J. Zhang, M. Guo, Y. Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang, Y. Kuang, M. Cao, F. Zheng, and X. Liang · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Gripper keypose and object pointflow as interfaces for bimanual robotic manipulation
Y. Yang, Z. Cai, Y. Tian, J. Zeng, and J. Pang · 2025
Closest in time.
Dexgraspvla: A vision-language-action framework towards general dexterous grasping
Y. Zhong, X. Huang, R. Li, C. Zhang, Y. Liang, Y. Yang, and Y. Chen · 2025
Closest in time.
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. M. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, O. H’enaff, J. Harmsen, A. Steiner, and X.-Q. Zhai · 2025
Closest in time.