Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have demonstrated significant potential in real-world robotic manipulation.
J. Wei, X. Wang et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , 2022
2022
Earlier work this paper cites.
Q. Vuong et al. , “Open x-embodiment: Robotic learning datasets and rt-x models,” in CoRL , 2023
2023
Earlier work this paper cites.
R. Firoozi et al. , “Foundation models in robotics: Applications, challenges, and the future,” IJRR , 2024
2024
Earlier work this paper cites.
M. Nakamoto et al. , “Steering your generalists: Improving robotic foundation models via value guidance,” in CoRL , 2024
2024
Earlier work this paper cites.
K. Black et al. , “Pi_0: A vision-language-action flow model for general robot control,” arXiv , 2024
2024
Earlier work this paper cites.
M. J. Kim et al. , “Openvla: An open-source vision-language-action model,” CoRL , 2024
2024
Earlier work this paper cites.
O. Mees et al. , “Octo: An open-source generalist robot policy,” in RSS , 2024
2024
Earlier work this paper cites.
M. Zawalski et al. , “Robotic control via embodied chain-of-thought reasoning,” in CoRL , 2024
2024
Earlier work this paper cites.
N. Ravi et al. , “Sam 2: Segment anything in images and videos,” arXiv , 2024
2024
Cited alongside, same era.
B. Yang et al. , “Diffusion-es: Gradient-free planning with diffusion for autonomous and instruction-guided driving,” in CVPR , 2024
2024
Cited alongside, same era.
Y. J. Ma et al. , “Eureka: Human-level reward design via coding large language models,” ICLR , 2024
2024
Cited alongside, same era.
A. Wagenmaker et al. , “Steering your diffusion policy with latent space reinforcement learning,” CoRL , 2025
2025
Cited alongside, same era.
Y. Wang et al. , “Inference-time policy steering through human interactions,” in IEEE ICRA , 2025
2025
Cited alongside, same era.
J. Wen et al. , “Diffusionvla: Scaling robot foundation models via unified diffusion and autoregression,” in ICML , 2025
2025
Closest in time.
Z. Li et al. , “Language-guided dexterous functional grasping by llm generated grasp functionality and synergy for humanoid manipulation,” IEEE T-ASE , 2025
2025
Closest in time.
Z. Li, J. Liu et al. , “Manidp: Manipulability-aware diffusion policy for posture-dependent bimanual manipulation,” IROS , 2025
2025
Closest in time.
J. Liu et al. , “Human–humanoid robots’ cross-embodiment behavior-skill transfer using decomposed adversarial learning from demonstration: Hotu, a human–humanoid robots’ skill transfer framework,” IEEE RAM , 2025
2025
Closest in time.
M. Dai, L. Liu et al. , “Rover: Robot reward model as test-time verifier for vision-language-action model,” arXiv , 2025
2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wu et al. , “From foresight to forethought: Vlm-in-the-loop policy steering via latent alignment,” in RSS , 2025
2025
Cited alongside, same era.
J. Bjorck et al. , “Gr00t n1: An open foundation model for generalist humanoid robots,” arXiv , 2025
2025
Cited alongside, same era.
S. Liu et al. , “RDT-1b: a diffusion foundation model for bimanual manipulation,” in ICLR , 2025
2025
Cited alongside, same era.
Closest in time.
C. Jose et al. , “Dinov2 meets text: A unified framework for image-and pixel-level vision-language alignment,” in CVPR , 2025
2025
Closest in time.
Y. Zhang et al. , “Diffusion models are evolutionary algorithms,” ICLR , 2025
2025
Closest in time.