Fetching the paper…
Reading the bibliography…
Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense,” in ICCV , 2017
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
T. Mu, Z. Ling, F. Xiang, D. Yang, X. Li, S. Tao, Z. Huang, Z. Jia, and H. Su, “ManiSkill: Generalizable manipulation skill benchmark with large-scale demonstrations,” in NeurIPS Datasets and Benchmarks , 2021
2021
Earlier work this paper cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022
2022
Earlier work this paper cites.
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine, “Bridge Data: Boosting generalization of robotic skills with cross-domain datasets,” in RSS , 2022
2022
Earlier work this paper cites.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al. , “Ego4D: Around the world in 3,000 hours of egocentric video,” in CVPR , 2022
2022
Earlier work this paper cites.
OpenAI, “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , 2023
2023
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn et al. , “RT-2: Vision-language-action models transfer web knowledge to robotic control,” in CoRL , 2023
2023
Earlier work this paper cites.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du et al. , “BridgeData v2: A dataset for robot learning at scale,” in CoRL , 2023
2023
Earlier work this paper cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion Policy: Visuomotor policy learning via action diffusion,” in RSS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al. , “PaLM-E: An embodied multimodal language model,” in ICML , 2023
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “RT-1: Robotics transformer for real-world control at scale,” in RSS , 2023
2023
Earlier work this paper cites.
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel, “Learning universal policies via text-guided video generation,” in NeurIPS , 2023
2023
Earlier work this paper cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV , 2023
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in ICCV , 2023
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al. , “InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in CVPR , 2024
2024
Cited alongside, same era.
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh, “Prismatic VLMs: Investigating the design space of visually-conditioned language models,” in ICML , 2024
2024
Cited alongside, same era.
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “DINOv2: Learning robust visual features without supervision,” TMLR , 2024
2024
Cited alongside, same era.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi et al. , “OpenVLA: An open-source vision-language-action model,” in CoRL , 2024
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter et al. , “ π \pi 0 {}_{\mbox{0}} : A vision-language-action flow model for general robot control,” in RSS , 2025
2025
Closest in time.
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “RDT-1B: a diffusion foundation model for bimanual manipulation,” in ICLR , 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis et al. , “DROID: A large-scale in-the-wild robot manipulation dataset,” in RSS , 2024
2024
Cited alongside, same era.
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan et al. , “Open X-Embodiment: Robotic learning datasets and RT-X models,” in ICRA , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
L. Wang, X. Chen, J. Zhao, and K. He, “Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers,” in NeurIPS , 2024
2024
Cited alongside, same era.
J. Yang, C. Glossop, A. Bhorkar, D. Shah, Q. Vuong, C. Finn, D. Sadigh, and S. Levine, “Pushing the limits of cross-embodiment learning for manipulation and navigation,” in RSS , 2024
2024
Cited alongside, same era.
R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine, “Scaling Cross-Embodied Learning: One policy for manipulation, navigation, locomotion and aviation,” in CoRL , 2024
2024
Cited alongside, same era.
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “MiniGPT-4: Enhancing vision-language understanding with advanced large language models,” in ICLR , 2024
2024
Cited alongside, same era.
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, H. Li, and T. Kong, “Vision-language foundation models as effective robot imitators,” in ICLR , 2024
2024
Cited alongside, same era.
Q. Bu, Y. Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li, “UniVLA: Learning to act anywhere with task-centric latent actions,” in RSS , 2025
2025
Closest in time.
2025
Closest in time.
L. Chen, C. Sima, K. Chitta, A. Loquercio, P. Luo, Y. Ma, and H. Li, “Intelligent Robot Manipulation Requires Self-Directed Learning,” Authorea Preprints , 2025
2025
Closest in time.
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao, “Data scaling laws in imitation learning for robotic manipulation,” in ICLR , 2025
2025
Closest in time.
J. Zheng, J. Li, D. Liu, Y. Zheng, Z. Wang, Z. Ou, Y. Liu, J. Liu, Y.-Q. Zhang, and X. Zhan, “Universal actions for enhanced embodied foundation models,” in CVPR , 2025
2025
Closest in time.
H. Li, Y. Cui, and D. Sadigh, “How to train your robots? the impact of demonstration modality on imitation learning,” in ICRA , 2025
2025
Closest in time.
J. Wen, Y. Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen et al. , “TinyVLA: Towards fast, data-efficient vision-language-action models for robotic manipulation,” RA-L , 2025
2025
Closest in time.
J. Wen, Y. Zhu, J. Li, Z. Tang, C. Shen, and F. Feng, “DexVLA: Vision-language model with plug-in diffusion expert for general robot control,” in CoRL , 2025
2025
Closest in time.
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin et al. , “Latent action pretraining from videos,” in ICLR , 2025
2025
Closest in time.
K. Wu, C. Hou, J. Liu, Z. Che, X. Ju, Z. Yang, M. Li, Y. Zhao, Z. Xu, G. Yang et al. , “RoboMIND: Benchmark on multi-embodiment intelligence normative data for robot manipulation,” in RSS , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
V. Saxena, M. Bronars, N. R. Arachchige, K. Wang, W. C. Shin, S. Nasiriany, A. Mandlekar, and D. Xu, “What matters in learning from large-scale datasets for robot manipulation,” in ICLR , 2025
2025
Closest in time.
Y. Mu, T. Chen, S. Peng, Z. Chen, Z. Gao, Y. Zou, L. Lin, Z. Xie, and P. Luo, “Robotwin: Dual-arm robot benchmark with generative digital twins (early version),” in ECCV , 2025
2025
Closest in time.
N. Masuya, S. Sakaino, and T. Tsuji, “Variable-frequency imitation learning for variable-speed motion,” in ICM , 2025
2025
Closest in time.
L. Guo, Z. Xue, Z. Xu, and H. Xu, “DemoSpeedup: Accelerating visuomotor policies via entropy-guided demonstration acceleration,” in CoRL , 2025
2025
Closest in time.
N. R. Arachchige, Z. Chen, W. Jung, W. C. Shin, R. Bansal, P. Barroso, Y. H. He, Y. C. Lin, B. Joffe, S. Kousik et al. , “SAIL: Faster-than-demonstration execution of imitation learning policies,” in CoRL , 2025
2025
Closest in time.
L. Wu, C. Yu, J. Ren, L. Chen, R. Huang, G. Gu, and H. Li, “FreeTacMan: Robot-free visuo-tactile data collection system for contact-rich manipulation,” in ICRA , 2026
2026
Closest in time.
Y. Pan, R. Qiao, L. Chen, K. Chitta, L. Pan, H. Mai, Q. Bu, C. Zheng, H. Zhao, P. Luo, and H. Li, “Agility Meets Stability: Versatile humanoid control with heterogeneous data,” in ICRA , 2026
2026
Closest in time.