Fetching the paper…
Reading the bibliography…
In this work, we present an approach to construct a video-based robot policy capable of reliably executing diverse tasks across different robots and environments from few video demonstrations without using any action annotations.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Human action recognition using factorized spatio-temporal convolutional networks
Lin Sun, Kui Jia, Dit-Yan Yeung, and Bertram E Shi · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep Visual Foresight for Planning Robot Motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression
Jangwon Lee and Michael S Ryoo · 2017
Earlier work this paper cites.
Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation
Peter R Florence, Lucas Manuelli, and Russ Tedrake · 2018
Earlier work this paper cites.
Learning Plannable Representations with Causal InfoGAN
Thanard Kurutach, Aviv Tamar, Ge Yang, Stuart J Russell, and Pieter Abbeel · 2018
Earlier work this paper cites.
An Algorithmic Perspective on Imitation Learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, Jan Peters, et al · 2018
Earlier work this paper cites.
Neural program synthesis from diverse demonstration videos
Shao-Hua Sun, Hyeonwoo Noh, Sriram Somasundaram, and Joseph Lim · 2018
Earlier work this paper cites.
Behavioral Cloning from Observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Earlier work this paper cites.
Goal-Conditioned Imitation Learning
Yiming Ding, Carlos Florensa, Pieter Abbeel, and Mariano Phielipp · 2019
Earlier work this paper cites.
Survey of Imitation Learning for Robotic Manipulation
Bin Fang, Shidong Jia, Di Guo, Muhua Xu, Shuhuan Wen, and Fuchun Sun · 2019
Earlier work this paper cites.
CompILE: Compositional Imitation Learning and Execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Earlier work this paper cites.
Third-person visual imitation learning via decoupled hierarchical controller
Pratyusha Sharma, Deepak Pathak, and Abhinav Gupta · 2019
Earlier work this paper cites.
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Earlier work this paper cites.
Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data
Priya Sundaresan, Jennifer Grannen, Brijen Thananjeyan, Ashwin Balakrishna, Michael Laskey, Kevin Stone, Joseph E Gonzalez, and Ken Goldberg · 2020
Earlier work this paper cites.
Learning Generalizable Robotic Reward Functions from ”In-The-Wild” Human Videos
Annie S Chen, Suraj Nair, and Chelsea Finn · 2021
Cited alongside, same era.
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Cited alongside, same era.
Generalizable Imitation Learning from Observation via Inferring Goal Proximity
Youngwoon Lee, Andrew Szot, Shao-Hua Sun, and Joseph J. Lim · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Reinforcement learning with videos: Combining offline observations with interaction
The Surprising Effectiveness of Representation Learning for Visual Imitation
Jyothish Pari, Nur Muhammad Shafiullah, Sridhar Pandian Arunachalam, and Lerrel Pinto · 2022
Later among the works it cites.
Progressive Distillation for Fast Sampling of Diffusion Models
Tim Salimans and Jonathan Ho · 2022
Later among the works it cites.
Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation
Anthony Simeonov, Yilun Du, Andrea Tagliasacchi, Joshua B Tenenbaum, Alberto Rodriguez, Pulkit Agrawal, and Vincent Sitzmann · 2022
Later among the works it cites.
Robotic Telekinesis: Learning a Robotic Hand Imitator by Watching Humans on Youtube
Aravind Sivakumar, Kenneth Shaw, and Deepak Pathak · 2022
Later among the works it cites.
GMFlow: Learning Optical Flow via Global Matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karl Schmeckpeper, Oleh Rybkin, Kostas Daniilidis, Sergey Levine, and Chelsea Finn · 2021
Cited alongside, same era.
Concept2Robot: Learning Manipulation Concepts from Instructions and Human Demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2021
Cited alongside, same era.
Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes
Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel, and Dieter Fox · 2021
Cited alongside, same era.
Human-to-robot imitation in the wild
Shikhar Bahl, Abhinav Gupta, and Deepak Pathak · 2022
Cited alongside, same era.
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos
Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Cited alongside, same era.
Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets
Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Daniilidis, Chelsea Finn, and Sergey Levine · 2022
Cited alongside, same era.
Lin Yen-Chen, Pete Florence, Jonathan T Barron, Tsung-Yi Lin, Alberto Rodriguez, and Phillip Isola · 2022
Later among the works it cites.
XIRL: Cross-embodiment inverse reinforcement learning
Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi · 2022
Later among the works it cites.
Local Neural Descriptor Fields: Locally Conditioned Object Representations for Manipulation
Ethan Chun, Yilun Du, Anthony Simeonov, Tomas Lozano-Perez, and Leslie Kaelbling · 2023
Closest in time.
Learning Universal Policies via Text-Guided Video Generation
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B Tenenbaum, Dale Schuurmans, and Pieter Abbeel · 2023
Closest in time.
Video Prediction Models as Rewards for Reinforcement Learning
Alejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain, Xue Bin Peng, Ken Goldberg, Youngwoon Lee, Danijar Hafner, and Pieter Abbeel · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-set Object Detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Closest in time.
Language Segment-Anything, 2023
Luca Medeiros · 2023
Closest in time.
Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation Learning
Hyunwoo Ryu, Jeong-Hoon Lee, Hong-in Lee, and Jongeun Choi · 2023
Closest in time.
SE(3)-Equivariant Relational Rearrangement with Neural Descriptor Fields
Anthony Simeonov, Yilun Du, Yen-Chen Lin, Alberto Rodriguez Garcia, Leslie Pack Kaelbling, Tomás Lozano-Pérez, and Pulkit Agrawal · 2023
Closest in time.
Diffusion Model-Augmented Behavioral Cloning
Hsiang-Chun Wang, Shang-Fu Chen, and Shao-Hua Sun · 2023
Closest in time.