Fetching the paper…
Reading the bibliography…
Developing efficient Vision-Language-Action (VLA) policies is crucial for practical robotics deployment, yet current approaches face prohibitive computational costs and resource requirements.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Hardware conditioned policies for multi-robot transfer learning
T. Chen, A. Murali, and A. Gupta · 2018
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Query-key normalization for transformers
A. Henry, P. R. Dachapally, S. Pawar, and Y. Chen · 2020
Earlier work this paper cites.
One policy to control them all: Shared modular policies for agent-agnostic control
W. Huang, I. Mordatch, and D. Pathak · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
X. Liu, C. Gong, and Q. Liu · 2022
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
M. S. Albergo and E. Vanden-Eijnden · 2022
Earlier work this paper cites.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, D. Sadigh, C. Finn, and S. Levine · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Earlier work this paper cites.
Open X-Embodiment: Robotic learning datasets and RT-X models
O. X.-E. Collaboration · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Robocat: A self-improving foundation agent for robotic manipulation
K. Bousmalis, G. Vezzani, D. Rao, C. Devin, A. X. Lee, M. Bauza, T. Davchev, Y. Zhou, A. Gupta, A. Raju, et al · 2023
Earlier work this paper cites.
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Vision-language foundation models as effective robot imitators
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, H. Li, and T. Kong · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Cited alongside, same era.
Polybot: Training one policy across robots while embracing variability
J. H. Yang, D. Sadigh, and C. Finn · 2023
Cited alongside, same era.
Goal conditioned imitation learning using score-based diffusion policies
M. Reuss, M. Li, X. Jia, and R. Lioutikov · 2023
Cited alongside, same era.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
Baku: An efficient transformer for multi-task policy learning
S. Haldar, Z. Peng, and L. Pinto · 2024
Later among the works it cites.
Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation
R. Doshi, H. R. Walke, O. Mees, S. Dasari, and S. Levine · 2024
Later among the works it cites.
The ingredients for robotic diffusion transformers
S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Cited alongside, same era.
Towards generalist robot policies: What matters in building vision-language-action models
X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu · 2024
Cited alongside, same era.
p i _ 0 pi\_0 : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Cited alongside, same era.
Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning, 2024
M. Reuss, J. Pari, P. Agrawal, and R. Lioutikov · 2024
Cited alongside, same era.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao · 2024
Cited alongside, same era.
Robot utility models: General policies for zero-shot deployment in new environments
H. Etukuru, N. Naka, Z. Hu, S. Lee, J. Mehu, A. Edsinger, C. Paxton, S. Chintala, L. Pinto, and N. M. M. Shafiullah · 2024
Cited alongside, same era.
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Later among the works it cites.
Unleashing large-scale video generative pre-training for visual robot manipulation
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang · 2024
Later among the works it cites.
Robouniview: Visual-language model with unified view representation for robotic manipulation
F. Liu, F. Yan, L. Zheng, Y. Huang, C. Feng, and L. Ma · 2024
Later among the works it cites.
Lerobot: State-of-the-art machine learning for real-world robotics in pytorch
R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf · 2024
Later among the works it cites.
A comprehensive user study on augmented reality-based data collection interfaces for robot learning
X. Jiang, P. Mattes, X. Jia, N. Schreiber, G. Neumann, and R. Lioutikov · 2024
Later among the works it cites.
Get-zero: Graph embodiment transformer for zero-shot embodiment generalization
A. Patel and S. Song · 2024
Later among the works it cites.
Towards diverse behaviors: A benchmark for imitation learning with human demonstrations
X. Jia, D. Blessing, X. Jiang, M. Reuss, A. Donat, R. Lioutikov, and G. Neumann · 2024
Later among the works it cites.
Actionflow: Equivariant, accurate, and efficient policies with spatially symmetric flow matching
N. Funk, J. Urain, J. Carvalho, V. Prasad, G. Chalvatzaki, and J. Peters · 2024
Later among the works it cites.
Affordance-based robot manipulation with flow matching
F. Zhang and M. Gienger · 2024
Later among the works it cites.
Riemannian flow matching policy for robot motion learning
M. Braun, N. Jaquier, L. Rozo, and T. Asfour · 2024
Later among the works it cites.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Smolvlm: Redefining small and efficient multimodal models
A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi, V. Srivastav, J. Lochner, H. Larcher, M. Morlon, L. Tunstall, L. von Werra, and T. Wolf · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Knowledge insulating vision-language-action models: Train fast, run fast, generalize better
D. Driess, J. T. Springenberg, B. Ichter, L. Yu, A. Li-Bell, K. Pertsch, A. Z. Ren, H. Walke, Q. Vuong, L. X. Shi, et al · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.