Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J.; and Schaal, S. 2007 · 2007
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Neural Ordinary Differential Equations
Chen, R. T. Q.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D. 2019 · 2019
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Kool, W.; van Hoof, H.; and Welling, M. 2019 · 2019
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Earlier work this paper cites.
Is Conditional Generative Modeling all you need for Decision Making?
Ajay, A.; Du, Y.; Gupta, A.; Tenenbaum, J. B.; Jaakkola, T. S.; and Agrawal, P. 2022 · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; et al. 2022 · 2022
Earlier work this paper cites.
Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling
Chen, H.; Lu, C.; Ying, C.; Su, H.; and Zhu, J. 2022 · 2022
Earlier work this paper cites.
Planning with Diffusion for Flexible Behavior Synthesis
Janner, M.; Du, Y.; Tenenbaum, J.; and Levine, S. 2022 · 2022
Earlier work this paper cites.
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
Wang, Z.; Hunt, J. J.; and Zhou, M. 2022 · 2022
Earlier work this paper cites.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
Black, K.; Nakamoto, M.; Atreya, P.; Walke, H.; Finn, C.; Kumar, A.; and Levine, S. 2023 · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023 · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al. 2023 · 2023
Cited alongside, same era.
Flow Matching for Generative Modeling
Lipman, Y.; Chen, R. T. Q.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2023 · 2023
Cited alongside, same era.
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reuss, M.; Yağmurlu, Ö. E.; Wenzel, F.; and Lioutikov, R. 2024 · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Team, O. M.; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; et al. 2024 · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, S.; Bai, H.; Lin, Z.; Pan, J.; Tong, P.; Zhou, Y.; Suhr, A.; Xie, S.; LeCun, Y.; Ma, Y.; et al. 2024 · 2024
Later among the works it cites.
Grape: Generalizing robot policy via preference alignment
Zhang, Z.; Zheng, K.; Chen, Z.; Jang, J.; Li, Y.; Wang, C.; Ding, M.; Fox, D.; and Yao, H. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2023 · 2023
Cited alongside, same era.
Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
Lu, C.; Chen, H.; Chen, J.; Su, H.; Li, C.; and Zhu, J. 2023 · 2023
Cited alongside, same era.
Guided flows for generative modeling and decision making
Zheng, Q.; Le, M.; Shaul, N.; Lipman, Y.; Grover, A.; and Chen, R. T. 2023 · 2023
Cited alongside, same era.
π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024 · 2024
Cited alongside, same era.
LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch
Cadene, R.; Alibert, S.; Soare, A.; Gallouedec, Q.; Zouitine, A.; Palma, S.; Kooijmans, P.; Aractingi, M.; Shukor, M.; Aubakirova, D.; Russi, M.; Capuano, F.; Pascale, C.; Choghari, J.; Moss, J.; and Wolf, T. 2024 · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024 · 2024
Cited alongside, same era.
Evaluating real-world robot manipulation policies in simulation
Li, X.; Hsu, K.; Gu, J.; Pertsch, K.; Mees, O.; Walke, H. R.; Fu, C.; Lunawat, I.; Sieh, I.; Kirmani, S.; et al. 2024 · 2024
Cited alongside, same era.
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Mark, M. S.; Gao, T.; Sampaio, G. G.; Srirama, M. K.; Sharma, A.; Finn, C.; and Kumar, A. 2024 · 2024
Cited alongside, same era.
Zhuang, Z.; Peng, D.; Liu, J.; Zhang, Z.; and Wang, D. 2024 · 2024
Later among the works it cites.
FDPP: Fine-tune Diffusion Policy with Human Preference
Chen, Y.; Jha, D. K.; Tomizuka, M.; and Romeres, D. 2025 · 2025
Closest in time.
Improving Vision-Language-Action Model with Online Reinforcement Learning
Guo, Y.; Zhang, J.; Chen, X.; Ji, X.; Wang, Y.-J.; Hu, Y.; and Chen, J. 2025 · 2025
Closest in time.
Dita: Scaling diffusion transformer for generalist vision-language-action policy
Hou, Z.; Zhang, T.; Xiong, Y.; Duan, H.; Pu, H.; Tong, R.; Zhao, C.; Zhu, X.; Qiao, Y.; Dai, J.; et al. 2025 · 2025
Closest in time.
Vla-rl: Towards masterful and general robotic manipulation with scalable reinforcement learning
Lu, G.; Guo, W.; Zhang, C.; Zhou, Y.; Jiang, H.; Gao, Z.; Tang, Y.; and Wang, Z. 2025 · 2025
Closest in time.
Interactive Post-Training for Vision-Language-Action Models
Tan, S.; Dou, K.; Zhao, Y.; and Krähenbühl, P. 2025 · 2025
Closest in time.
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
Zhang, H.; Zhuang, Z.; Zhao, H.; Ding, P.; Lu, H.; and Wang, D. 2025 · 2025
Closest in time.
Energy-Weighted Flow Matching for Offline Reinforcement Learning
Zhang, S.; Zhang, W.; and Gu, Q. 2025 · 2025
Closest in time.
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
Zhao, H.; Song, W.; Wang, D.; Tong, X.; Ding, P.; Cheng, X.; and Ge, Z. 2025 · 2025
Closest in time.