Fetching the paper…
Reading the bibliography…
This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as pouring, wiping, and mixing--purely by imitating AI-generated videos, without requiring any physical demonstrations or robot-specific training.
Retargetting motion to new characters
Michael Gleicher · 1998
Earlier work this paper cites.
A flexible new technique for camera calibration
Zhengyou Zhang · 2000
Earlier work this paper cites.
Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman · 2003
Earlier work this paper cites.
Task model of lower body motion for a biped humanoid robot to imitate human dances
Shinichiro Nakaoka, Atsushi Nakazawa, Fumio Kanehiro, Kenji Kaneko, Mitsuharu Morisawa, and Katsushi Ikeuchi · 2005
Earlier work this paper cites.
Ep n p: An accurate o (n) solution to the p n p problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua · 2009
Earlier work this paper cites.
Learning and generalization of complex tasks from unstructured demonstrations
Scott Niekum, Sarah Osentoski, George Konidaris, and Andrew G Barto · 2012
Earlier work this paper cites.
Online human walking imitation in task and joint space based on quadratic programming
Kai Hu, Christian Ott, and Dongheui Lee · 2014
Earlier work this paper cites.
A tutorial on task-parameterized movement learning and retrieval
Sylvain Calinon · 2016
Earlier work this paper cites.
Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot
Scott Kuindersma, Robin Deits, Maurice Fallon, Andrés Valenzuela, Hongkai Dai, Frank Permenter, Twan Koolen, Pat Marion, and Russ Tedrake · 2016
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
Chelsea Finn, Tianhe Yu, T. Zhang, P. Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Dense object nets: Learning dense visual object descriptors by and for robotic manipulation
Peter R Florence, Lucas Manuelli, and Russ Tedrake · 2018
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, et al · 2018
Earlier work this paper cites.
Zero-shot visual imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A Efros, and Trevor Darrell · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain · 2018
Earlier work this paper cites.
Multiple interactions made easy (mime): Large scale demonstrations data for imitation
Pratyusha Sharma, Lekha Mohan, Lerrel Pinto, and Abhinav Gupta · 2018
Earlier work this paper cites.
Deep object pose estimation for semantic robotic grasping of household objects
Jonathan Tremblay, Thang To, Balakumar Sundaralingam, Yu Xiang, Dieter Fox, and Stan Birchfield · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
Tianhe Yu, Chelsea Finn, Annie Xie, Sudeep Dasari, Tianhao Zhang, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Pix2pose: Pixel-wise coordinate regression of objects for 6d pose estimation
Kiru Park, Timothy Patten, and Markus Vincze · 2019
Earlier work this paper cites.
A multimode teleoperation framework for humanoid loco-manipulation: An application for the icub robot
Luigi Penco, Nicola Scianca, Valerio Modugno, Leonardo Lanari, Giuseppe Oriolo, and Serena Ivaldi · 2019
Earlier work this paper cites.
Third-person visual imitation learning via decoupled hierarchical controller
Pratyusha Sharma, Deepak Pathak, and Abhinav Gupta · 2019
Earlier work this paper cites.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
Laura Smith, Nikita Dhawan, Marvin Zhang, Pieter Abbeel, and Sergey Levine · 2019
Earlier work this paper cites.
Flowcontrol: Optical flow based visual servoing
Max Argus, Lukas Hermann, Jon Long, and Thomas Brox · 2020
Earlier work this paper cites.
Reconstruct locally, localize globally: A model free method for object pose estimation
Ming Cai and Ian Reid · 2020
Earlier work this paper cites.
Semantic visual navigation by watching youtube videos
Matthew Chang, Arjun Gupta, and Saurabh Gupta · 2020
Earlier work this paper cites.
Nonparametric motion retargeting for humanoid robots on shared latent space
Sungjoon Choi, Matthew KXJ Pan, and Joohyung Kim · 2020
Earlier work this paper cites.
Pvn3d: A deep point-wise 3d keypoints voting network for 6dof pose estimation
Yisheng He, Wei Sun, Haibin Huang, Jianran Liu, Haoqiang Fan, and Jian Sun · 2020
Earlier work this paper cites.
Cosypose: Consistent multi-view multi-object 6d pose estimation
Yann Labbé, Justin Carpentier, Mathieu Aubry, and Josef Sivic · 2020
Earlier work this paper cites.
Latentfusion: End-to-end differentiable reconstruction and rendering for unseen object pose estimation
Keunhong Park, Arsalan Mousavian, Yu Xiang, and Dieter Fox · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Earlier work this paper cites.
Transformers for one-shot visual imitation
Sudeep Dasari and Abhinav Gupta · 2021
Earlier work this paper cites.
Ffb6d: A full flow bidirectional fusion network for 6d pose estimation
Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, and Jian Sun · 2021
Earlier work this paper cites.
Dynamic movement primitive based motion retargeting for dual-arm sign language motions
Yuwei Liang, Weijie Li, Yue Wang, Rong Xiong, Yichao Mao, and Jiafan Zhang · 2021
Earlier work this paper cites.
Amp: Adversarial motion priors for stylized physics-based character control
Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Human-to-robot imitation in the wild
Shikhar Bahl, Abhinav Gupta, and Deepak Pathak · 2022
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Earlier work this paper cites.
Megapose: 6d pose estimation of novel objects via render & compare
Yann Labbé, Lucas Manuelli, Arsalan Mousavian, Stephen Tyree, Stan Birchfield, Jonathan Tremblay, Justin Carpentier, Mathieu Aubry, Dieter Fox, and Josef Sivic · 2022
Earlier work this paper cites.
Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images
Yuan Liu, Yilin Wen, Sida Peng, Cheng Lin, Xiaoxiao Long, Taku Komura, and Wenping Wang · 2022
Earlier work this paper cites.
Learning to imitate object interactions from internet videos
Austin Patel, Andrew Wang, Ilija Radosavovic, and Jitendra Malik · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang · 2022
Earlier work this paper cites.
Osop: A multi-stage one shot object pose estimation framework
Ivan Shugurov, Fu Li, Benjamin Busam, and Slobodan Ilic · 2022
Cited alongside, same era.
Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube
Aravind Sivakumar, Kenneth Shaw, and Deepak Pathak · 2022
Cited alongside, same era.
Onepose: One-shot object pose estimation without cad models
Jiaming Sun, Zihao Wang, Siyu Zhang, Xingyi He, Hongcheng Zhao, Guofeng Zhang, and Xiaowei Zhou · 2022
Cited alongside, same era.
Demonstrate once, imitate immediately (dome): Learning visual servoing for one-shot imitation learning
Eugene Valassakis, Georgios Papagiannis, Norman Di Palo, and Edward Johns · 2022
Cited alongside, same era.
Gmflow: Learning optical flow via global matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao · 2022
Cited alongside, same era.
Humanplus: Humanoid shadowing and imitation from humans
Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, and Chelsea Finn · 2024
Later among the works it cites.
Flip: Flow-centric generative planning for general-purpose manipulation tasks
Chongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai, and Lin Shao · 2024
Later among the works it cites.
Learning human-to-humanoid real-time whole-body teleoperation
Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi · 2024
Later among the works it cites.
Spot: Se (3) pose trajectory diffusion for object-centric manipulation
Cheng-Chun Hsu, Bowen Wen, Jie Xu, Yashraj Narang, Xiaolong Wang, Yuke Zhu, Joydeep Biswas, and Stan Birchfield · 2024
Later among the works it cites.
Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xirl: Cross-embodiment inverse reinforcement learning
Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Compositional foundation models for hierarchical planning
Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li, Abhi Gupta, Tommi Jaakkola, Josh Tenenbaum, Leslie Kaelbling, Akash Srivastava, and Pulkit Agrawal · 2023
Cited alongside, same era.
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak · 2023
Cited alongside, same era.
Zero-shot robot manipulation from passive human videos
Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar · 2023
Cited alongside, same era.
Learning video-conditioned policies for unseen manipulation tasks
Elliot Chane-Sane, Cordelia Schmid, and Ivan Laptev · 2023
Cited alongside, same era.
Matthew Chang, Theophile Gervet, Mukul Khanna, Sriram Yenamandra, Dhruv Shah, So Yeon Min, Kavit Shah, Chris Paxton, Saurabh Gupta, Dhruv Batra, Roozbeh Mottaghi, Jitendra Malik, and Devendra Singh Chaplot · 2023
Cited alongside, same era.
Yuanchen Ju, Kaizhe Hu, Guowei Zhang, Gu Zhang, Mingrun Jiang, and Huazhe Xu · 2024
Later among the works it cites.
Cotracker3: Simpler and better point tracking by pseudo-labelling real videos
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2024
Later among the works it cites.
Egomimic: Scaling imitation learning via egocentric video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu · 2024
Later among the works it cites.
Video depth without video models, 2024
Bingxin Ke, Dominik Narnhofer, Shengyu Huang, Lei Ke, Torben Peters, Katerina Fragkiadaki, Anton Obukhov, and Konrad Schindler · 2024
Later among the works it cites.
Robot see robot do: Imitating articulated object manipulation with monocular 4d reconstruction
Justin Kerr, Chung Min Kim, Mingxuan Wu, Brent Yi, Qianqian Wang, Ken Goldberg, and Angjoo Kanazawa · 2024
Later among the works it cites.
Garfield: Group anything with radiance fields
Chung Min Kim, Mingxuan Wu, Justin Kerr, Ken Goldberg, Matthew Tancik, and Angjoo Kanazawa · 2024
Later among the works it cites.
Kinematic motion retargeting for contact-rich anthropomorphic manipulations
Arjun S Lakshmipathy, Jessica K Hodgins, and Nancy S Pollard · 2024
Later among the works it cites.
Dreamitate: Real-world visuomotor policy learning via video generation
Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, and Carl Vondrick · 2024
Later among the works it cites.
Gigapose: Fast and robust novel object pose estimation via one correspondence
Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, and Vincent Lepetit · 2024
Later among the works it cites.
Foundpose: Unseen object pose estimation with foundation features
Evin Pınar Örnek, Yann Labbé, Bugra Tekin, Lingni Ma, Cem Keskin, Christian Forster, and Tomáš Hodaň · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer · 2024
Later among the works it cites.
Hrp: Human affordances for robotic pre-training
Mohan Kumar Srirama, Sudeep Dasari, Shikhar Bahl, and Abhinav Gupta · 2024
Later among the works it cites.
Video creation by demonstration
Yihong Sun, Hao Zhou, Liangzhe Yuan, Jennifer J Sun, Yandong Li, Xuhui Jia, Hartwig Adam, Bharath Hariharan, Long Zhao, and Ting Liu · 2024
Later among the works it cites.
Robotap: Tracking arbitrary points for few-shot visual imitation
Mel Vecerik, Carl Doersch, Yi Yang, Todor Davchev, Yusuf Aytar, Guangyao Zhou, Raia Hadsell, Lourdes Agapito, and Jon Scholz · 2024
Later among the works it cites.
One-shot video imitation via parameterized symbolic abstraction graphs
Jianren Wang, Kangni Liu, Dingkun Guo, Xian Zhou, and Christopher G Atkeson · 2024
Later among the works it cites.
Exploitation-guided exploration for semantic embodied navigation
Justin Wasserman, Girish Chowdhary, Abhinav Gupta, and Unnat Jain · 2024
Later among the works it cites.
One-shot transfer of long-horizon extrinsic manipulation through contact retargeting
Albert Wu, Ruocheng Wang, Sirui Chen, Clemens Eppner, and C Karen Liu · 2024
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
Annie Xie, Lisa Lee, Ted Xiao, and Chelsea Finn · 2024
Later among the works it cites.
Flow as the cross-domain manipulation interface
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song · 2024
Later among the works it cites.
Latent action pretraining from videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al · 2024
Later among the works it cites.
General flow as foundation affordance for scalable robot learning
Chengbo Yuan, Chuan Wen, Tong Zhang, and Yang Gao · 2024
Later among the works it cites.
World-consistent video diffusion with explicit 3d modeling
Qihang Zhang, Shuangfei Zhai, Miguel Angel Bautista, Kevin Miao, Alexander Toshev, Joshua Susskind, and Jiatao Gu · 2024
Later among the works it cites.
Mitigating the human-robot domain discrepancy in visual pre-training for robotic manipulation
Jiaming Zhou, Teli Ma, Kun-Yu Lin, Zifan Wang, Ronghe Qiu, and Junwei Liang · 2024
Later among the works it cites.
Densematcher: Learning 3d semantic correspondence for category-level manipulation from a single demo
Junzhe Zhu, Yuanchen Ju, Junyi Zhang, Muhan Wang, Zhecheng Yuan, Kaizhe Hu, and Huazhe Xu · 2024
Later among the works it cites.
Nil: No-data imitation learning by leveraging pre-trained video diffusion models
Mert Albaba, Chenhao Li, Markos Diomataris, Omid Taheri, Andreas Krause, and Michael Black · 2025
Closest in time.
Adaworld: Learning adaptable world models with latent actions
Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, and Chuang Gan · 2025
Closest in time.
T2vphysbench: A first-principles benchmark for physical consistency in text-to-video generation
Xuyang Guo, Jiayan Huo, Zhenmei Shi, Zhao Song, Jiahao Zhang, and Jiale Zhao · 2025
Closest in time.
Phantom: Training robots without robots using only human videos
Marion Lepert, Jiaying Fang, and Jeannette Bohg · 2025
Closest in time.
Do generative video models learn physical principles from watching videos?
Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos · 2025
Closest in time.
A real-to-sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards
Shivansh Patel, Xinchen Yin, Wenlong Huang, Shubham Garg, Hooshang Nayyeri, Li Fei-Fei, Svetlana Lazebnik, and Yunzhu Li · 2025
Closest in time.
6d object pose tracking in internet videos for robotic manipulation
Georgy Ponimatkin, Martin Cífka, Tomáš Souček, Médéric Fourmy, Yann Labbé, Vladimir Petrik, and Josef Sivic · 2025
Closest in time.
Motion tracks: A unified representation for human-robot transfer in few-shot imitation learning
Juntao Ren, Priya Sundaresan, Dorsa Sadigh, Sanjiban Choudhury, and Jeannette Bohg · 2025
Closest in time.
Zeromimic: Distilling robotic manipulation skills from web videos
Junyao Shi, Zhuolun Zhao, Tianyou Wang, Ian Pedroza, Amy Luo, Jie Wang, Jason Ma, and Dinesh Jayaraman · 2025
Closest in time.
Vlipp: Towards physically plausible video generation with vision and language informed physical prior
Xindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, et al · 2025
Closest in time.
Tesseract: Learning 4d embodied world models
Haoyu Zhen, Qiao Sun, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, and Chuang Gan · 2025
Closest in time.
You only teach once: Learn one-shot bimanual robotic manipulation from video demonstrations
Huayi Zhou, Ruixiang Wang, Yunxin Tai, Yueci Deng, Guiliang Liu, and Kui Jia · 2025
Closest in time.