Fetching the paper…
Reading the bibliography…
Pre-trained vision-language-action (VLA) models offer a promising foundation for generalist robot policies, but often produce brittle behaviors or unsafe failures when deployed zero-shot in out-of-distribution scenarios.
“Mastering the game of Go with deep neural networks and tree search”
David Silver et al · 2016
Earlier work this paper cites.
“Mastering the game of go without human knowledge”
David Silver et al · 2017
Earlier work this paper cites.
“A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play”
David Silver et al · 2018
Earlier work this paper cites.
“robosuite: A modular simulation framework and benchmark for robot learning”
Yuke Zhu et al · 2020
Earlier work this paper cites.
“Dream to Control: Learning Behaviors by Latent Imagination”
Danijar Hafner, Timothy Lillicrap, Jimmy Ba and Mohammad Norouzi · 2020
Earlier work this paper cites.
“Mastering atari, go, chess and shogi by planning with a learned model”
Julian Schrittwieser et al · 2020
Earlier work this paper cites.
“Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning”
Viktor Makoviychuk et al · 2021
Earlier work this paper cites.
“Mastering Atari with Discrete World Models”
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi and Jimmy Ba · 2021
Earlier work this paper cites.
“Learning and planning in complex action spaces”
Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt and David Silver · 2021
Earlier work this paper cites.
“Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”
Jason Wei et al · 2022
Earlier work this paper cites.
“Temporal Difference Learning for Model Predictive Control”
Nicklas Hansen, Hao Su and Xiaolong Wang · 2022
Earlier work this paper cites.
“Do as i can, not as i say: Grounding language in robotic affordances”
Michael Ahn et al · 2022
Earlier work this paper cites.
“Monte Carlo tree search: A review of recent modifications and applications”
Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki and Jacek Mańdziuk · 2023
Earlier work this paper cites.
“LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning”
Bo Liu et al · 2023
Earlier work this paper cites.
“RT-1: Robotics Transformer for Real-World Control at Scale”
Anthony Brohan et al · 2023
Earlier work this paper cites.
“RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control”
Brianna Zitkovich et al · 2023
Earlier work this paper cites.
“Perceiver-actor: A multi-task transformer for robotic manipulation”
Mohit Shridhar, Lucas Manuelli and Dieter Fox · 2023
Cited alongside, same era.
“PaLM-E: An Embodied Multimodal Language Model”
Danny Driess et al · 2023
Cited alongside, same era.
“Tree of thoughts: Deliberate problem solving with large language models”
Shunyu Yao et al · 2023
Cited alongside, same era.
“Self-Consistency Improves Chain of Thought Reasoning in Language Models”
Xuezhi Wang et al · 2023
Cited alongside, same era.
“Reasoning with Language Model is Planning with World Model”
Shibo Hao et al · 2023
Cited alongside, same era.
“Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training”
Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang and Jun Wang · 2023
Cited alongside, same era.
“Video Language Planning”
Yilun Du et al · 2024
Later among the works it cites.
“Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming” 8th IFAC Conference on Nonlinear Model Predictive Control NMPC 2024
Dimitri. Bertsekas · 2024
Later among the works it cites.
“TD-MPC2: Scalable, Robust World Models for Continuous Control”
Nicklas Hansen, Hao Su and Xiaolong Wang · 2024
Later among the works it cites.
“LGMCTS: Language-Guided Monte-Carlo Tree Search for Executable Semantic Object Rearrangement”
Haonan Chang et al · 2024
Later among the works it cites.
“Robotic Control via Embodied Chain-of-Thought Reasoning”
Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn and Sergey Levine · 2024
Later among the works it cites.
“Vision Language Models are In-Context Value Learners”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks”
Murtaza Dalal, Tarun Chiruvolu, Devendra Chaplot and Ruslan Salakhutdinov · 2023
Cited alongside, same era.
“Zhao, Tony Z and Kumar, Vikash and Levine, Sergey and Finn, Chelsea”
Learning fine-grained-cost hardware · 2023
Cited alongside, same era.
“Octo: An Open-Source Generalist Robot Policy”
Octo Model Team et al · 2024
Cited alongside, same era.
“DROID: A large-scale in-the-wild robot manipulation dataset”
Alexander Khazatsky et al · 2024
Cited alongside, same era.
“Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0”
Abby O’Neill et al · 2024
Cited alongside, same era.
“OpenVLA: An Open-Source Vision-Language-Action Model”
Moo Kim et al · 2024
Cited alongside, same era.
Yecheng Ma et al · 2024
Later among the works it cites.
“Octo: An Open-Source Generalist Robot Policy”
Octo Model Team et al · 2024
Later among the works it cites.
“Paligemma: A versatile 3b vlm for transfer”
Lucas Beyer et al · 2024
Later among the works it cites.
“RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies”
Pranav Atreya et al · 2025
Closest in time.
“Fast: Efficient action tokenization for vision-language-action models”
Karl Pertsch et al · 2025
Closest in time.
“Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Though”
Violet Xiang et al · 2025
Closest in time.
“Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning”
Daya Guo et al · 2025
Closest in time.
“Test-time Computing: from System-1 Thinking to System-2 Thinking”
Yixin Ji et al · 2025
Closest in time.
“From system 1 to system 2: A survey of reasoning large language models”
Zhong-Zhi Li et al · 2025
Closest in time.
“Reflective planning: Vision-Language Models for multi-stage long-horizon robotic manipulation”
Yunhai Feng, Jiaming Han, Zhuoran Yang, Xiangyu Yue, Sergey Levine and Jianlan Luo · 2025
Closest in time.
“Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models”
Lucy Shi et al · 2025
Closest in time.