Fetching the paper…
Reading the bibliography…
Embodied Chain-of-Thought (ECoT) reasoning enhances vision-language-action (VLA) models by improving performance and interpretability through intermediate reasoning steps.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” 2021
2021
Earlier work this paper cites.
M. Ahn and A. e. a. Brohan, “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,” in Conference on Robot Learning , 2022, pp. 287–318
2022
Earlier work this paper cites.
A. Zeng and M. A. et al., “Socratic models: Composing zero-shot multimodal reasoning with language,” 2022
2022
Earlier work this paper cites.
G.-I. Yu and J. S. J. et al., “Orca: A distributed serving system for Transformer-Based generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . Carlsbad, CA: USENIX Association, July 2022, pp. 521–538
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” in Conference on Neural Information Processing Systems (NeurIPS) , 2022
2022
Earlier work this paper cites.
S. K. et al., “QuaRL: Quantization for fast and environmentally sustainable reinforcement learning,” Transactions on Machine Learning Research , 2022
2022
Earlier work this paper cites.
H. Walke and K. e. a. Black, “Bridgedata v2: A dataset for robot learning at scale,” in Conference on Robot Learning (CoRL) , 2023
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Earlier work this paper cites.
J. Wei and X. W. et al., “Chain-of-thought prompting elicits reasoning in large language models,” 2023
2023
Earlier work this paper cites.
O. Mees, J. Borja-Diaz, and W. Burgard, “Grounding language with visual affordances over unstructured data,” 2023
2023
Earlier work this paper cites.
Y. L. et al., “Fast inference from transformers via speculative decoding,” 2023
2023
Earlier work this paper cites.
S. Kim and K. M. et al., “Speculative decoding with big little decoder,” 2023
2023
Earlier work this paper cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and Efficient Post-Training Quantization for Large Language Models.” in International Conference on Machine Learning (ICML) , 2023, pp. 38 087–38 099
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.-C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositional reasoning with large language models,” 2023
2023
Earlier work this paper cites.
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
K. Black and N. B. et al., “ π 0 \pi_{0} : A vision-language-action flow model for general robot control,” 2024
2024
Cited alongside, same era.
C.-P. Huang, Y.-H. Wu, M.-H. Chen, Y.-C. F. Wang, and F.-E. Yang, “Thinkact: Vision-language-action reasoning via reinforced visual latent planning,” 2025
2025
Closest in time.
J. Lee and J. D. et al., “Molmoact: Action reasoning models that can reason in space,” 2025
2025
Closest in time.
S. Wang, R. Yu, Z. Yuan, C. Yu, F. Gao, Y. Wang, and D. F. Wong, “Spec-vla: Speculative decoding for vision-language-action models with relaxed acceptance,” 2025
2025
Closest in time.
X. Tan, Y. Yang, P. Ye, J. Zheng, B. Bai, X. Wang, J. Hao, and T. Chen, “Think twice, act once: Token-aware compression and action reuse for efficient inference in vision-language-action models,” 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
A. O’Neill and A. R. et al., “Open x-embodiment: Robotic learning datasets and rt-x models,” 2024
2024
Cited alongside, same era.
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh, “Prismatic vlms: Investigating the design space of visually-conditioned language models,” 2024
2024
Cited alongside, same era.
X. Ning and Z. L. et al., “Skeleton-of-thought: Prompting llms for efficient parallel generation,” 2024
2024
Cited alongside, same era.
Q. Sun, P. Hong, T. D. Pala, V. Toh, U.-X. Tan, D. Ghosal, and S. Poria, “Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning,” 2024
2024
Cited alongside, same era.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for llm compression and acceleration,” 2024
2024
Cited alongside, same era.
NVIDIA, “Gr00t n1: An open foundation model for generalist humanoid robots,” 2025
2025
Cited alongside, same era.
W. Song, J. Chen, P. Ding, H. Zhao, W. Zhao, Z. Zhong, Z. Ge, J. Ma, and H. Li, “Accelerating vision-language-action model integrated with action chunking via parallel decoding,” 2025
2025
Cited alongside, same era.
2025
Closest in time.
huggingface, “Bitsandbytes,” 2025. [Online]. Available: https://huggingface.co/docs/bitsandbytes/main/en/index
2025
Closest in time.
Nvidia, “Tensorrt-llm.” 2025
2025
Closest in time.
DeepSeek-AI, D. Guo, and D. Y. et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025
2025
Closest in time.
G. Xu, P. Jin, H. Li, Y. Song, L. Sun, and L. Yuan, “Llava-cot: Let vision language models reason step-by-step,” 2025
2025
Closest in time.
M. S. et al., “Smolvla: A vision-language-action model for affordable and efficient robotics,” 2025
2025
Closest in time.
NVIDIA, :, and J. B. et al., “Gr00t n1: An open foundation model for generalist humanoid robots,” 2025
2025
Closest in time.
2025
Closest in time.
huggingface, “Bitsandbytes,” 2025
2025
Closest in time.
A. Khazatsky and K. P. et al., “Droid: A large-scale in-the-wild robot manipulation dataset,” 2025
2025
Closest in time.
Z. Zhou and P. A. et al., “Autoeval: Autonomous evaluation of generalist robot manipulation policies in the real world,” 2025
2025
Closest in time.