Fetching the paper…
Reading the bibliography…
Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter. arxiv 2019
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Earlier work this paper cites.
Sapien: A simulated part-based interactive environment
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, et al · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
H.-S. Fang, C. Wang, M. Gou, and C. Lu · 2020
Earlier work this paper cites.
Learning to see before learning to act: Visual pre-training for manipulation
L. Yen-Chen, A. Zeng, S. Song, P. Isola, and T.-Y. Lin · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching
A. Zeng, S. Song, K.-T. Yu, E. Donlon, F. R. Hogan, M. Bauza, D. Ma, O. Taylor, M. Liu, E. Romo, et al · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Compositional visual generation with composable diffusion models
N. Liu, S. Li, Y. Du, A. Torralba, and J. B. Tenenbaum · 2022
Earlier work this paper cites.
Robocook: Long-horizon elasto-plastic object manipulation with diverse tools
H. Shi, H. Xu, S. Clarke, Y. Li, and J. Wu · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao · 2023
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, et al · 2023
Earlier work this paper cites.
Tidybot: Personalized robot assistance with large language models
J. Wu, R. Antonova, A. Kan, M. Lepert, A. Zeng, S. Song, J. Bohg, S. Rusinkiewicz, and T. Funkhouser · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Earlier work this paper cites.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Earlier work this paper cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Earlier work this paper cites.
Fleet policy learning via weight merging and an application to robotic tool-use
L. Wang, K. Zhang, A. Zhou, M. Simchowitz, and R. Tedrake · 2023
Cited alongside, same era.
Learning visuotactile skills with two multifingered hands
T. Lin, Y. Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik · 2024
Cited alongside, same era.
Data scaling laws in imitation learning for robotic manipulation
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2024
Cited alongside, same era.
Learning manipulation skills through robot chain-of-thought with sparse failure guidance
K. Zhang, Z.-H. Yin, W. Ye, and Y. Gao · 2024
Cited alongside, same era.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Diffusion forcing: Next-token prediction meets full-sequence diffusion
Y. Du, M. Simchowitz, R. Tedrake, V. Sitzmann, B. Chen, and D. M. Monso · 2024
Later among the works it cites.
Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning
Y. Wang, Y. Zhang, M. Huo, R. Tian, X. Zhang, Y. Xie, C. Xu, P. Ji, W. Zhan, M. Ding, et al · 2024
Later among the works it cites.
Consistency policy: Accelerated visuomotor policies via consistency distillation
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg · 2024
Later among the works it cites.
The ingredients for robotic diffusion transformers
S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine · 2024
Later among the works it cites.
Data scaling laws in imitation learning for robotic manipulation, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov · 2024
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Cited alongside, same era.
3d-vla: A 3d vision-language-action generative world model
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Cited alongside, same era.
Humanplus: Humanoid shadowing and imitation from humans
Z. Fu, Q. Zhao, Q. Wu, G. Wetzstein, and C. Finn · 2024
Cited alongside, same era.
Umi on legs: Making manipulation policies mobile with manipulation-centric whole-body controllers
H. Ha, Y. Gao, Z. Fu, J. Tan, and S. Song · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Cited alongside, same era.
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2024
Later among the works it cites.
Diffusion policy policy optimization
A. Z. Ren, J. Lidard, L. L. Ankile, A. Simeonov, P. Agrawal, A. Majumdar, B. Burchfiel, H. Dai, and M. Simchowitz · 2024
Later among the works it cites.
Inference-time policy steering through human interactions
Y. Wang, L. Wang, Y. Du, B. Sundaralingam, X. Yang, Y.-W. Chao, C. Perez-D’Arpino, D. Fox, and J. Shah · 2024
Later among the works it cites.
Video prediction policy: A generalist robot policy with predictive visual representations
Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Later among the works it cites.
Dnact: Diffusion guided multi-task 3d policy learning
G. Yan, Y.-H. Wu, and X. Wang · 2024
Later among the works it cites.
Diffusion-vla: Scaling robot foundation models via unified diffusion and autoregression
J. Wen, M. Zhu, Y. Zhu, Z. Tang, J. Li, Z. Zhou, C. Li, X. Liu, Y. Peng, C. Shen, et al · 2024
Later among the works it cites.
Discrete policy: Learning disentangled action space for multi-task robotic manipulation
K. Wu, Y. Zhu, J. Li, J. Wen, N. Liu, Z. Xu, Q. Qiu, and J. Tang · 2024
Later among the works it cites.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al · 2025
Closest in time.
pi0.5: a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Improving vision-language-action model with online reinforcement learning
Y. Guo, J. Zhang, X. Chen, X. Ji, Y.-J. Wang, Y. Hu, and J. Chen · 2025
Closest in time.