Fetching the paper…
Reading the bibliography…
Most Vision-Language-Action (VLA) systems integrate a Vision-Language Model (VLM) for semantic reasoning with an action expert generating continuous action signals, yet both typically run at a single unified frequency.
“Machine thinking, fast and slow”
Jean-François Bonnefon and Iyad Rahwan · 2020
Earlier work this paper cites.
“Autoregressive image generation using residual quantization”
Doyup Lee et al · 2022
Earlier work this paper cites.
“RT-1: Robotics Transformer for Real-World Control at Scale”, 2023
Anthony et al · 2023
Earlier work this paper cites.
“RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control”
Brianna Zitkovich et al · 2023
Earlier work this paper cites.
“Octo: An Open-Source Generalist Robot Policy”, 2024
Octo Team et al · 2024
Earlier work this paper cites.
“BridgeData V2: A Dataset for Robot Learning at Scale”, 2024
Homer Walke et al · 2024
Earlier work this paper cites.
“Openvla: An open-source vision-language-action model”
Moo Kim et al · 2024
Earlier work this paper cites.
“Rdt-1b: a diffusion foundation model for bimanual manipulation”
Songming Liu et al · 2024
Earlier work this paper cites.
“A dual process vla: Efficient robotic manipulation leveraging vlm”
ByungOk Han, Jaehong Kim and Jinhyeok Jang · 2024
Earlier work this paper cites.
“Towards synergistic, generalized, and efficient dual-system for robotic manipulation”
Qingwen Bu et al · 2024
Earlier work this paper cites.
“From llms to actions: Latent codes as bridges in hierarchical robot control”
Yide Shentu, Philipp Wu, Aravind Rajeswaran and Pieter Abbeel · 2024
Earlier work this paper cites.
“Paligemma: A versatile 3b vlm for transfer”
Lucas Beyer et al · 2024
Earlier work this paper cites.
“Cllms: Consistency large language models”
Siqi Kou et al · 2024
Earlier work this paper cites.
Kevin Black et al · 2025
Earlier work this paper cites.
“GR00T N1: An Open Foundation Model for Generalist Humanoid Robots”, 2025
NVIDIA et al · 2025
Cited alongside, same era.
“RynnEC: Bringing MLLMs into Embodied World”, 2025
Ronghao Dang et al · 2025
Cited alongside, same era.
“ π 0.6 ∗ \pi^{*}_{0.6} : a VLA That Learns From Experience”, 2025
Physical Intelligence et al · 2025
Cited alongside, same era.
“GEN-0: Embodied Foundation Models That Scale with Physical Interaction” https://generalistai.com/blog/preview-uqlxvb-bb.html
Generalist Team · 2025
Cited alongside, same era.
“Helix: A Vision-Language-Action Model for Generalist Humanoid Control” https://www.figure.ai/news/helix
Figure AI · 2025
Cited alongside, same era.
“ π 0.5 \pi_{0.5} : a Vision-Language-Action Model with Open-World Generalization”, 2025
Physical Intelligence et al · 2025
Closest in time.
“Igniting vlms toward the embodied space”
Andy Zhai et al · 2025
Closest in time.
“Gr00t n1: An open foundation model for generalist humanoid robots”
Johan Bjorck et al · 2025
Closest in time.
“A0: An affordance-aware hierarchical model for general robotic manipulation”
Rongtao Xu et al · 2025
Closest in time.
“Eo-1: Interleaved vision-text-action pretraining for general robot control”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Hume: Introducing System-2 Thinking in Visual-Language-Action Model”, 2025
Haoming Song et al · 2025
Cited alongside, same era.
Hao Chen et al · 2025
Cited alongside, same era.
Can Cui et al · 2025
Cited alongside, same era.
“RDT2: Enabling Zero-Shot Cross-Embodiment Generalization by Scaling Up UMI Data”, 2025
RDT Team · 2025
Cited alongside, same era.
“Cot-vla: Visual chain-of-thought reasoning for vision-language-action models”
Qingqing Zhao et al · 2025
Cited alongside, same era.
“Fast: Efficient action tokenization for vision-language-action models”
Karl Pertsch et al · 2025
Cited alongside, same era.
“Universal actions for enhanced embodied foundation models”
Jinliang Zheng et al · 2025
Cited alongside, same era.
Delin Qu et al · 2025
Closest in time.
“Interactive Post-Training for Vision-Language-Action Models”
Shuhan Tan, Kairan Dou, Yue Zhao and Philipp Krähenbühl · 2025
Closest in time.
“Vla-rl: Towards masterful and general robotic manipulation with scalable reinforcement learning”
Guanxing Lu et al · 2025
Closest in time.
“ π RL \pi_{\texttt{RL}} : Online RL Fine-tuning for Flow-based Vision-Language-Action Models”, 2025
Kang Chen et al · 2025
Closest in time.
“ π 0.6 ∗ \pi^{*}_{0.6} : a VLA That Learns From Experience”, 2025
Physical Intelligence et al · 2025
Closest in time.
“Scaling diffusion policy in transformer to 1 billion parameters for robotic manipulation”
Minjie Zhu et al · 2025
Closest in time.
“Galaxea open-world dataset and g0 dual-system vla model”
Tao Jiang et al · 2025
Closest in time.
“Qwen2. 5-vl technical report”
Shuai Bai et al · 2025
Closest in time.
“Towards Human-level Intelligence via Human-like Whole-Body Manipulation”, 2025
Guang Gao et al · 2025
Closest in time.