Fetching the paper…
Reading the bibliography…
Learning robot control policies from demonstrations typically requires action-labeled expert data, which is expensive to collect through teleoperation.
“D4RL: Datasets for Deep Data-Driven Reinforcement Learning”, 2020
Justin Fu et al · 2004
Earlier work this paper cites.
“MuJoCo: A physics engine for model-based control”
Emanuel Todorov, Tom Erez and Yuval Tassa · 2012
Earlier work this paper cites.
“Neural discrete representation learning”
Aaron Van, Oriol Vinyals and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
“Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation”
YuXuan Liu et al · 2018
Earlier work this paper cites.
“Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement”
Chao Yang et al · 2019
Earlier work this paper cites.
“State Alignment-based Imitation Learning”
Fangchen Liu et al · 2020
Earlier work this paper cites.
“Learning latent plans from play”
Corey Lynch et al · 2020
Earlier work this paper cites.
“Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning”
Tianhe Yu et al · 2020
Earlier work this paper cites.
“Playable Video Generation”
Willi Menapace et al · 2021
Earlier work this paper cites.
“Is space-time attention all you need for video understanding?”
Gedas Bertasius, Heng Wang and Lorenzo Torresani · 2021
Earlier work this paper cites.
“Ego4D: Around the World in 3,000 Hours of Egocentric Video”
Kristen Grauman et al · 2022
Earlier work this paper cites.
“Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100”
Dima Damen et al · 2022
Earlier work this paper cites.
“Video pretraining (vpt): Learning to act by watching unlabeled online videos”
Bowen Baker et al · 2022
Earlier work this paper cites.
“Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks”
Oier Mees et al · 2022
Earlier work this paper cites.
“Robotube: Learning household manipulation from human videos with simulated twin environments”
Haoyu Xiong et al · 2023
Cited alongside, same era.
“Open x-embodiment: Robotic learning datasets and rt-x models”
Open X-Embodiment Collaboration · 2023
Cited alongside, same era.
“Any-point Trajectory Modeling for Policy Learning”, 2023
Chuan Wen et al · 2023
Cited alongside, same era.
“Learning Universal Policies via Text-Guided Video Generation”
Yilun Du et al · 2023
Cited alongside, same era.
“Learning to Act without Actions”
Dominik Schmidt and Minqi Jiang · 2023
Cited alongside, same era.
“Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware”
“HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval”, 2025
Matthew Hong et al · 2025
Closest in time.
“Phantom: Training Robots Without Robots Using Only Human Videos”
Marion Lepert, Jiaying Fang and Jeannette Bohg · 2025
Closest in time.
“Masquerade: Learning from In-the-wild Human Videos using Data-Editing”
Marion Lepert, Jiaying Fang and Jeannette Bohg · 2025
Closest in time.
“Object-Centric Latent Action Learning”
Albina Klepach et al · 2025
Closest in time.
“Learning to Act Anywhere with Task-centric Latent Actions”
Qingwen Bu et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tony. Zhao et al · 2023
Cited alongside, same era.
“Bridgedata v2: A dataset for robot learning at scale”
Homer Walke et al · 2023
Cited alongside, same era.
“ π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control”
Kevin Black et al · 2024
Cited alongside, same era.
“Learning to Act from Actionless Videos through Dense Correspondences”
Po-Chen Ko et al · 2024
Cited alongside, same era.
“Predictive inverse dynamics models are scalable learners for robotic manipulation”
Yang Tian et al · 2024
Cited alongside, same era.
“Genie: Generative Interactive Environments”
Jake Bruce et al · 2024
Cited alongside, same era.
“DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control”
Zichen Cui et al · 2024
Cited alongside, same era.
“villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models”
Xiaoyu Chen et al · 2025
Closest in time.
“Latent Action Pretraining Through World Modeling”, 2025
Bahey Tharwat et al · 2025
Closest in time.
“Latent Action Learning Requires Supervision in the Presence of Distractors”
Alexander Nikulin et al · 2025
Closest in time.
“CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning”, 2025
Jiange Yang et al · 2025
Closest in time.
“Learning Latent Action World Models In The Wild”, 2026
Quentin Garrido et al · 2026
Closest in time.
Fanqi Lin et al · 2026
Closest in time.
“World Action Models are Zero-shot Policies”, 2026
Seonghyeon Ye et al · 2026
Closest in time.
“DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos”
Shenyuan Gao et al · 2026
Closest in time.