Fetching the paper…
Reading the bibliography…
We present Im2Flow2Act, a scalable learning framework that enables robots to acquire real-world manipulation skills without the need of real-world robot training data.
System identification and control using genetic algorithms
K. Kristinsson and G. Dumont · 1992
Earlier work this paper cites.
Policy distillation, 2016
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2016
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Y. Liu, A. Gupta, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Training deep networks with synthetic data: Bridging the reality gap by domain randomization
J. Tremblay, A. Prakash, D. Acuna, M. Brophy, V. Jampani, C. Anil, T. To, E. Cameracci, S. Boochoon, and S. Birchfield · 2018
Earlier work this paper cites.
Implicit 3d orientation learning for 6d object detection from rgb images
M. Sundermeyer, Z.-C. Marton, M. Durner, M. Brucker, and R. Triebel · 2018
Earlier work this paper cites.
Third-person visual imitation learning via decoupled hierarchical controller
P. Sharma, D. Pathak, and A. Gupta · 2019
Earlier work this paper cites.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine · 2019
Earlier work this paper cites.
Real-world robotic perception and control using synthetic data
J. Tobin · 2019
Earlier work this paper cites.
Domain randomization and pyramid consistency: Simulation-to-real generalization without accessing target domain data
X. Yue, Y. Zhang, S. Zhao, A. Sangiovanni-Vincentelli, K. Keutzer, and B. Gong · 2019
Earlier work this paper cites.
Continual reinforcement learning deployed in real-life using policy distillation and sim2real transfer
R. Traoré, H. Caselles-Dupré, T. Lesort, T. Sun, N. Díaz-Rodríguez, and D. Filliat · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Learning predictive models from observation and interaction
K. Schmeckpeper, A. Xie, O. Rybkin, S. Tian, K. Daniilidis, S. Levine, and C. Finn · 2020
Earlier work this paper cites.
Sim2real transfer for reinforcement learning without dynamics randomization
M. Kaspar, J. D. M. Osorio, and J. Bock · 2020
Earlier work this paper cites.
Learning Generalizable Robotic Reward Functions from “In-The-Wild” Human Videos
A. S. Chen, S. Nair, and C. Finn · 2021
Earlier work this paper cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2021
Earlier work this paper cites.
Reinforcement learning with videos: Combining offline observations with interaction
K. Schmeckpeper, O. Rybkin, K. Daniilidis, S. Levine, and C. Finn · 2021
Earlier work this paper cites.
Learning by watching: Physical imitation of manipulation skills from human videos
H. Xiong, Q. Li, Y.-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis, 2021
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Y. Zhu and A. Joshi · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Cited alongside, same era.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Cited alongside, same era.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Cited alongside, same era.
Mimicplay: Long-horizon imitation learning by watching human play
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Later among the works it cites.
Xskill: Cross embodiment skill discovery
M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Later among the works it cites.
Tapir: Tracking any point with per-frame initialization and temporal refinement
C. Doersch, Y. Yang, M. Vecerik, D. Gokay, A. Gupta, Y. Aytar, J. Carreira, and A. Zisserman · 2023
Later among the works it cites.
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. C. Burchfiel, and S. Song · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Flowbot3d: Learning 3d articulation flow to manipulate articulated objects
B. Eisner, H. Zhang, and D. Held · 2022
Cited alongside, same era.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Cited alongside, same era.
Xirl: Cross-embodiment inverse reinforcement learning
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2022
Cited alongside, same era.
Human hands as probes for interactive object understanding
M. Goyal, S. Modi, R. Goyal, and S. Gupta · 2022
Cited alongside, same era.
Joint hand motion and interaction hotspots prediction from egocentric videos
S. Liu, S. Tripathi, S. Majumdar, and X. Wang · 2022
Cited alongside, same era.
Giving robots a hand: Broadening generalization via hand-centric human video demonstrations
M. J. Kim, J. Wu, and C. Finn · 2022
Cited alongside, same era.
Attention is all you need, 2023
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2023
Later among the works it cites.
Close the optical sensing domain gap by physics-grounded active stereo sensor simulation
X. Zhang, R. Chen, A. Li, F. Xiang, Y. Qin, J. Gu, Z. Ling, M. Liu, P. Zeng, S. Han, et al · 2023
Later among the works it cites.
Dexpoint: Generalizable point cloud reinforcement learning for sim-to-real dexterous manipulation
Y. Qin, B. Huang, Z.-H. Yin, H. Su, and X. Wang · 2023
Later among the works it cites.
Sim2real transfer learning for point cloud segmentation: An industrial application case on autonomous disassembly
C. Wu, X. Bi, J. Pfrommer, A. Cebulla, S. Mangold, and J. Beyerer · 2023
Later among the works it cites.
Y. Huang, J. Yuan, C. Kim, P. Pradhan, B. Chen, L. Fuxin, and T. Hermans · 2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Later among the works it cites.
Vision-based manipulation from single human video with open-world object graphs, 2024
Y. Zhu, A. Lim, P. Stone, and Y. Zhu · 2024
Closest in time.
General flow as foundation affordance for scalable robot learning
C. Yuan, C. Wen, T. Zhang, and Y. Gao · 2024
Closest in time.
Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation, 2024
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani · 2024
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, T. Wu, J. Vakil, et al · 2024
Closest in time.
A comprehensive survey of cross-domain policy transfer for embodied agents, 2024
H. Niu, J. Hu, G. Zhou, and X. Zhan · 2024
Closest in time.
Towards generalist robot learning from internet video: A survey, 2024
R. McCarthy, D. C. H. Tan, D. Schmidt, F. Acero, N. Herr, Y. Du, T. G. Thuruthel, and Z. Li · 2024
Closest in time.
Peac: Unsupervised pre-training for cross-embodiment reinforcement learning, 2024
C. Ying, Z. Hao, X. Zhou, X. Xu, H. Su, X. Zhang, and J. Zhu · 2024
Closest in time.
Dexcap: Scalable and portable mocap data collection system for dexterous manipulation
C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu · 2024
Closest in time.
Mirage: Cross-embodiment zero-shot policy transfer with cross-painting, 2024
L. Y. Chen, K. Hari, K. Dharmarajan, C. Xu, Q. Vuong, and K. Goldberg · 2024
Closest in time.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Tapvid-3d: A benchmark for tracking any point in 3d, 2024
S. Koppula, I. Rocco, Y. Yang, J. Heyward, J. Carreira, A. Zisserman, G. Brostow, and C. Doersch · 2024
Closest in time.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Closest in time.