Fetching the paper…
Reading the bibliography…
In this paper, we present DiffusionVLA, a novel framework that seamlessly combines the autoregression model with the diffusion model for learning visuomotor policy.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Earlier work this paper cites.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
Gu, J., Kirmani, S., Wohlhart, P., Lu, Y., Arenas, M. G., Rao, K., Yu, W., Fu, C., Gopalakrishnan, K., Xu, Z., et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., et al · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Earlier work this paper cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
Rt-h: Action hierarchies using language
Belkhale, S., Ding, T., Xiao, T., Sermanet, P., Vuong, Q., Tompson, J., Chebotar, Y., Dwibedi, D., and Sadigh, D · 2024
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Shi, L. X., Tanner, J., Vuong, Q., Walling, A., Wang, H., and Zhilinsky, U · 2024
Cited alongside, same era.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reuss, M., Yağmurlu, Ö. E., Wenzel, F., and Lioutikov, R · 2024
Closest in time.
Yell at your robot: Improving on-the-fly from language corrections
Shi, L. X., Hu, Z., Zhao, T. Z., Sharma, A., Pertsch, K., Luo, J., Levine, S., and Finn, C · 2024
Closest in time.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
Team, C · 2024
Closest in time.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bu, Q., Li, H., Chen, L., Cai, J., Zeng, J., Cui, H., Yao, M., and Qiao, Y · 2024
Cited alongside, same era.
The ingredients for robotic diffusion transformers
Dasari, S., Mees, O., Zhao, S., Srirama, M. K., and Levine, S · 2024
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Fu, Z., Zhao, T. Z., and Finn, C · 2024
Cited alongside, same era.
Seed-x: Multimodal models with unified multi-granularity comprehension and generation
Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., and Shan, Y · 2024
Cited alongside, same era.
An embodied generalist agent in 3d world
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.-C., Jia, B., and Huang, S · 2024
Cited alongside, same era.
3d diffuser actor: Policy diffusion with 3d scene representations
Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
Khazatsky, A., Pertsch, K., Nair, S., Balakrishna, A., Dasari, S., Karamcheti, S., Nasiriany, S., Srirama, M. K., Chen, L. Y., Ellis, K., et al · 2024
Cited alongside, same era.
Data scaling laws in imitation learning for robotic manipulation, 2024
Lin, F., Hu, Y., Sheng, P., Wen, C., You, J., and Gao, Y · 2024
Cited alongside, same era.
Wen, J., Zhu, Y., Li, J., Zhu, M., Wu, K., Xu, Z., Cheng, R., Shen, C., Peng, Y., Feng, F., et al · 2024
Closest in time.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Closest in time.
Dnact: Diffusion guided multi-task 3d policy learning
Yan, G., Wu, Y.-H., and Wang, X · 2024
Closest in time.
Learning to manipulate anywhere: A visual generalizable framework for reinforcement learning
Yuan, Z., Wei, T., Cheng, S., Zhang, G., Chen, Y., and Xu, H · 2024
Closest in time.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Ze, Y., Zhang, G., Zhang, K., Hu, C., Wang, M., and Xu, H · 2024
Closest in time.
Grape: Generalizing robot policy via preference alignment
Zhang, Z., Zheng, K., Chen, Z., Jang, J., Li, Y., Wang, C., Ding, M., Fox, D., and Yao, H · 2024
Closest in time.
Monoformer: One transformer for both diffusion and autoregression
Zhao, C., Song, Y., Wang, W., Feng, H., Ding, E., Sun, Y., Xiao, X., and Wang, J · 2024
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Closest in time.
Scaling diffusion policy in transformer to 1 billion parameters for robotic manipulation
Zhu, M., Zhu, Y., Li, J., Wen, J., Xu, Z., Liu, N., Cheng, R., Shen, C., Peng, Y., Feng, F., et al · 2024
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., and Levine, S · 2025
Closest in time.
Universal actions for enhanced embodied foundation models
Zheng, J., Li, J., Liu, D., Zheng, Y., Wang, Z., Ou, Z., Liu, Y., Liu, J., Zhang, Y.-Q., and Zhan, X · 2025
Closest in time.