Fetching the paper…
Reading the bibliography…
Learning whole-body mobile manipulation via imitation is essential for generalizing robotic skills to diverse environments and complex tasks.
“ALVINN: An Autonomous Land Vehicle in a Neural Network”
Dean Pomerleau · 1988
Earlier work this paper cites.
“PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation”
Charles Qi, Hao Su, Kaichun Mo and Leonidas Guibas · 2016
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“4d spatio-temporal convnets: Minkowski convolutional neural networks”
Christopher Choy, JunYoung Gwak and Silvio Savarese · 2019
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Denoising diffusion probabilistic models”
J. Ho, A. Jain and P. Abbeel · 2020
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu et al · 2021
Earlier work this paper cites.
“BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning”
Jang, E. et al · 2021
Earlier work this paper cites.
“Denoising diffusion implicit models”
J. Song, C. Meng and S. Ermon · 2021
Earlier work this paper cites.
“Implicit behavioral cloning”
P. Florence et al · 2022
Earlier work this paper cites.
“Planning with Diffusion for Flexible Behavior Synthesis”, 2022
Michael Janner, Yilun Du, Joshua. Tenenbaum and Sergey Levine · 2022
Earlier work this paper cites.
“Is Conditional Generative Modeling all you need for Decision-Making?”, 2023
Anurag Ajay et al · 2023
Earlier work this paper cites.
“Structured Denoising Diffusion Models in Discrete State-Spaces”, 2023
Jacob Austin et al · 2023
Earlier work this paper cites.
“Diffusion Policy: Visuomotor Policy Learning via Action Diffusion”
Chi, C. et al · 2023
Earlier work this paper cites.
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Earlier work this paper cites.
“Visual Instruction Tuning”, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Earlier work this paper cites.
“DINOv2: Learning Robust Visual Features without Supervision”
Maxime Oquab, Timothée Darcet and Moutakanni et al · 2023
Earlier work this paper cites.
“Sigmoid Loss for Language Image Pre-Training”, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov and Lucas Beyer · 2023
Cited alongside, same era.
“Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware”
Zhao, T. Z. et al · 2023
Cited alongside, same era.
“ π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control”, 2024
Kevin Black et al · 2024
Cited alongside, same era.
“Open X-Embodiment: Robotic Learning Datasets and RT-X Models”
Open-Embodiment Collaboration et al · 2024
Cited alongside, same era.
“In-Context Imitation Learning via Next-Token Prediction”
Letian Fu et al · 2024
Cited alongside, same era.
“AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation”, 2025
Sixiang Chen et al · 2025
Closest in time.
“Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better”, 2025
Danny Driess et al · 2025
Closest in time.
“AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons”
Hongjie Fang and Chenxi et al · 2025
Closest in time.
“VITA: Vision-to-Action Flow Matching Policy”, 2025
Dechen Gao et al · 2025
Closest in time.
“Towards Human-level Intelligence via Human-like Whole-Body Manipulation”, 2025
Guang Gao and Jianan et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhefei Gong et al · 2024
Cited alongside, same era.
Yueru Jia et al · 2024
Cited alongside, same era.
“OpenVLA: An Open-Source Vision-Language-Action Model”
Moo Kim, Karl Pertsch and Karamcheti et al · 2024
Cited alongside, same era.
“Octo: An Open-Source Generalist Robot Policy”
Octo Model Team · 2024
Cited alongside, same era.
“RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective”
Chenxi Wang, Hongjie Fang, Hao-Shu Fang and Cewu Lu · 2024
Cited alongside, same era.
“CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation”
Shangning Xia, Hongjie Fang, Hao-Shu Fang and Cewu Lu · 2024
Cited alongside, same era.
“Generalizable Humanoid Manipulation with 3D Diffusion Policies”
Yanjie Ze et al · 2024
Cited alongside, same era.
Closest in time.
“D-AR: Diffusion via Autoregressive Models”, 2025
Ziteng Gao and Mike Shou · 2025
Closest in time.
“ π 0.5 \pi_{0.5} : a Vision-Language-Action Model with Open-World Generalization”, 2025
Physical Intelligence and Kevin et al · 2025
Closest in time.
“BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities”
Yunfan Jiang et al · 2025
Closest in time.
“H 3 DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning”, 2025
Yiyang Lu et al · 2025
Closest in time.
“FAST: Efficient Action Tokenization for Vision-Language-Action Models”, 2025
Karl Pertsch et al · 2025
Closest in time.
“Dense Policy: Bidirectional Autoregressive Learning of Actions”
Yue Su et al · 2025
Closest in time.
“Motion Before Action: Diffusing Object Motion as Manipulation Condition”
Yue Su et al · 2025
Closest in time.
“UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation”, 2025
Yihe Tang et al · 2025
Closest in time.
Han Xue et al · 2025
Closest in time.
“FP3: A 3D Foundation Policy for Robotic Manipulation”, 2025
Rujia Yang, Geng Chen, Chuan Wen and Yang Gao · 2025
Closest in time.
Zhecheng Yuan et al · 2025
Closest in time.