Fetching the paper…
Reading the bibliography…
Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training.
Approximated centroidal voronoi diagrams for uniform polygonal mesh coarsening
S. Valette and J.-M. Chassery · 2004
Earlier work this paper cites.
Acronym: A large-scale grasp dataset based on simulation, 2020
C. Eppner, A. Mousavian, and D. Fox · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Using simulation and domain adaptation to improve efficiency of deep robotic grasping, 2017
K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige, S. Levine, and V. Vanhoucke · 2017
Earlier work this paper cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
Gpu-accelerated robotic simulation for distributed reinforcement learning, 2018
J. Liang, V. Makoviychuk, A. Handa, N. Chentanez, M. Macklin, and D. Fox · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation, 2018
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Earlier work this paper cites.
On evaluation of embodied navigation agents, 2018
P. Anderson, A. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, and A. R. Zamir · 2018
Earlier work this paper cites.
6-dof graspnet: Variational grasp generation for object manipulation
A. Mousavian, C. Eppner, and D. Fox · 2019
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
H.-S. Fang, C. Wang, M. Gou, and C. Lu · 2020
Earlier work this paper cites.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
S. Song, A. Zeng, J. Lee, and T. Funkhouser · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
URL https://github.com/google-deepmind/envlogger
google-deepmind/envlogger, Jan. 2025a · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Earlier work this paper cites.
Chatgpt: Jan 17 version
OpenAI · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, et al · 2023
Earlier work this paper cites.
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu · 2023
Earlier work this paper cites.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking, 2023
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Earlier work this paper cites.
Pali-x: On scaling up a multilingual vision and language model
X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, Y. Tay, et al · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Earlier work this paper cites.
Mimicgen: A data generation system for scalable robot learning using human demonstrations, 2023
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Earlier work this paper cites.
Genaug: Retargeting behaviors to unseen situations via generative augmentation, 2023
Z. Chen, S. Kiami, A. Gupta, and V. Kumar · 2023
Earlier work this paper cites.
Scaling robot learning with semantically imagined experience, 2023
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, D. M, J. Peralta, B. Ichter, K. Hausman, and F. Xia · 2023
Earlier work this paper cites.
Deep learning approaches to grasp synthesis: A review
R. Newbury, M. Gu, L. Chumbley, A. Mousavian, C. Eppner, J. Leitner, J. Bohg, A. Morales, T. Asfour, D. Kragic, et al · 2023
Cited alongside, same era.
H. Geng, S. Wei, C. Deng, B. Shen, H. Wang, and L. Guibas · 2023
Cited alongside, same era.
Grasp-anything: Large-scale grasp dataset from foundation models, 2023
A. D. Vuong, M. N. Vu, H. Le, B. Huang, B. Huynh, T. Vo, A. Kugi, and A. Nguyen · 2023
Cited alongside, same era.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, S. Kirmani, B. Zitkovich, F. Xia, et al · 2023
Cited alongside, same era.
Graspgpt: Leveraging semantic knowledge from a large language model for task-oriented grasping
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation, 2024
J. Wen, Y. Zhu, J. Li, M. Zhu, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, Y. Peng, F. Feng, and J. Tang · 2024
Later among the works it cites.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Later among the works it cites.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, et al · 2024
Later among the works it cites.
Latent action pretraining from videos
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Tang, D. Huang, W. Ge, W. Liu, and H. Zhang · 2023
Cited alongside, same era.
Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes
Y. Lu, Y. Fan, B. Deng, F. Liu, Y. Li, and S. Wang · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi · 2023
Cited alongside, same era.
Curobo: Parallelized collision-free robot motion generation
B. Sundaralingam, S. K. S. Hari, A. Fishman, C. Garrett, K. Van Wyk, V. Blukis, A. Millane, H. Oleynikova, A. Handa, F. Ramos, et al · 2023
Cited alongside, same era.
Orbit: A unified simulation framework for interactive robot learning environments
M. Mittal, C. Yu, Q. Yu, J. Liu, N. Rudin, D. Hoeller, J. L. Yuan, R. Singh, Y. Guo, H. Mazhar, A. Mandlekar, B. Babich, G. State, M. Hutter, and A. Garg · 2023
Cited alongside, same era.
Imitating task and motion planning with visuomotor transformers
M. Dalal, A. Mandlekar, C. Garrett, A. Handa, R. Salakhutdinov, and D. Fox · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani · 2024
Later among the works it cites.
Spatiotemporal predictive pre-training for robotic motor control
J. Yang, B. Liu, J. Fu, B. Pan, G. Wu, and L. Wang · 2024
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation, 2024
Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang · 2024
Later among the works it cites.
Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment, 2024
C. Garrett, A. Mandlekar, B. Wen, and D. Fox · 2024
Later among the works it cites.
D3roma: Disparity diffusion-based depth sensing for material-agnostic robotic manipulation
S. Wei, H. Geng, J. Chen, C. Deng, C. Wenbo, C. Zhao, X. Fang, L. Guibas, and H. Wang · 2024
Later among the works it cites.
Efficient end-to-end detection of 6-dof grasps for robotic bin picking, 2024
Y. Liu, A. Qualmann, Z. Yu, M. Gabriel, P. Schillinger, M. Spies, N. A. Vien, and A. Geiger · 2024
Later among the works it cites.
Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Later among the works it cites.
Open6dor: Benchmarking open-instruction 6-dof object rearrangement and a vlm-based approach
Y. Ding, H. Geng, C. Xu, X. Fang, J. Zhang, S. Wei, Q. Dai, Z. Zhang, and H. Wang · 2024
Later among the works it cites.
Bodex: Scalable and efficient robotic dexterous grasp synthesis using bilevel optimization
J. Chen, Y. Ke, and H. Wang · 2024
Later among the works it cites.
Data scaling laws in imitation learning for robotic manipulation, 2024
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2024
Later among the works it cites.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu, et al · 2024
Later among the works it cites.
Paligemma: A versatile 3b vlm for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al · 2024
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang · 2024
Later among the works it cites.
Bridgedata v2: A dataset for robot learning at scale, 2024
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine · 2024
Later among the works it cites.
Serl: A software suite for sample-efficient robotic reinforcement learning, 2024
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots, 2025
NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, et al · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization, 2025
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky · 2025
Closest in time.
Z. Jiang, Y. Xie, K. Lin, Z. Xu, W. Wan, A. Mandlekar, L. Fan, and Y. Zhu · 2025
Closest in time.
Novel demonstration generation with gaussian splatting enables robust one-shot manipulation, 2025
S. Yang, W. Yu, J. Zeng, J. Lv, K. Ren, C. Lu, D. Lin, and J. Pang · 2025
Closest in time.
Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning, 2025
Z. Xue, S. Deng, Z. Chen, Y. Wang, Z. Yuan, and H. Xu · 2025
Closest in time.
Sim-and-real co-training: A simple recipe for vision-based robotic manipulation, 2025
A. Maddukuri, Z. Jiang, L. Y. Chen, S. Nasiriany, Y. Xie, Y. Fang, W. Huang, Z. Wang, Z. Xu, N. Chernyadev, S. Reed, K. Goldberg, A. Mandlekar, L. Fan, and Y. Zhu · 2025
Closest in time.
Y. Chen, B. Xiao, and H. Wang · 2025
Closest in time.
AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, S. Jiang, Y. Jiang, C. Jing, H. Li, J. Li, C. Liu, Y. Liu, Y. Lu, J. Luo, P. Luo, Y. Mu, Y. Niu, Y. Pan, J. Pang, Y. Qiao, G. Ren, C. Ruan, J. Shan, Y. Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu, J. Yan, C. Yang, L. Yang, S. Yang, M. Yao, J. Zeng, C. Zhang, Q. Zhang, B. Zhao, C. Zhao, J. Zhao, and J. Zhu · 2025
Closest in time.