Fetching the paper…
Reading the bibliography…
Robot learning of manipulation skills is hindered by the scarcity of diverse, unbiased datasets.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
C. Dima, M. Hebert, and A. Stentz, “Enabling learning from large datasets: applying active learning to mobile robotics,” in IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA ’04. 2004 , vol. 1, 2004, pp. 108–114 Vol.1
2004
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems , vol. 57, no. 5, pp. 469–483, 2009
2009
Earlier work this paper cites.
B. Settles, “Active learning literature survey,” 2009
2009
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: a large video database for human motion recognition,” in 2011 International conference on computer vision . IEEE, 2011, pp. 2556–2563
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
H. S. Koppula, R. Gupta, and A. Saxena, “Learning human activities and object affordances from rgb-d videos,” The International journal of robotics research , vol. 32, no. 8, pp. 951–970, 2013
2013
Earlier work this paper cites.
A. Pieropan, C. H. Ek, and H. Kjellström, “Functional object descriptors for human activity modeling,” in 2013 IEEE International Conference on Robotics and Automation . IEEE, 2013, pp. 1282–1289
2013
Earlier work this paper cites.
W. Zhang, M. Zhu, and K. G. Derpanis, “From actemes to action: A strongly-supervised representation for detailed action understanding,” in 2013 IEEE International Conference on Computer Vision , 2013, pp. 2248–2255
2013
Earlier work this paper cites.
H. S. Koppula and A. Saxena, “Physically grounded spatio-temporal object affordances,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part III 13 . Springer, 2014, pp. 831–847
2014
Earlier work this paper cites.
D. I. Kim and G. S. Sukhatme, “Semantic labeling of 3d point clouds with object affordance for robot manipulation,” in 2014 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2014, pp. 5578–5584
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, “2d human pose estimation: New benchmark and state of the art analysis,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2014
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, pp. 211–252, 2015
2015
Earlier work this paper cites.
X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2794–2802
2015
Earlier work this paper cites.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” in International conference on machine learning . PMLR, 2015, pp. 843–852
2015
Earlier work this paper cites.
H. S. Koppula and A. Saxena, “Anticipating human activities using object affordances for reactive robotic response,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 1, pp. 14–29, 2015
2015
Earlier work this paper cites.
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “SMPL: A skinned multi-person linear model,” ACM Trans. Graphics (Proc. SIGGRAPH Asia) , vol. 34, no. 6, pp. 248:1–248:16, Oct. 2015
2015
Earlier work this paper cites.
Y. Yang, Y. Li, C. Fermuller, and Y. Aloimonos, “Robot learning manipulation action plans by” watching” unconstrained videos from the world wide web,” in Proceedings of the AAAI conference on artificial intelligence , vol. 29, no. 1, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Flynn, I. Neulander, J. Philbin, and N. Snavely, “Deepstereo: Learning to predict new views from the world’s imagery,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 5515–5524
2016
Earlier work this paper cites.
R. Garg, V. K. Bg, G. Carneiro, and I. Reid, “Unsupervised cnn for single view depth estimation: Geometry to the rescue,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14 . Springer, 2016, pp. 740–756
2016
Earlier work this paper cites.
M. Cai, K. M. Kitani, and Y. Sato, “Understanding hand-object manipulation with grasp types and object attributes.” in Robotics: Science and Systems , vol. 3. Ann Arbor, Michigan;, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine, “One-shot visual imitation learning via meta-learning,” in Conference on robot learning . PMLR, 2017, pp. 357–368
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1851–1858
2017
Earlier work this paper cites.
C. Godard, O. Mac Aodha, and G. J. Brostow, “Unsupervised monocular depth estimation with left-right consistency,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 270–279
2017
Earlier work this paper cites.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5842–5850
2017
Earlier work this paper cites.
J. Lee and M. S. Ryoo, “Learning robot activities from first-person human videos using convolutional future regression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2017, pp. 1–2
2017
Earlier work this paper cites.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1125–1134
2017
Earlier work this paper cites.
J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, and E. Shechtman, “Toward multimodal image-to-image translation,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
M.-Y. Liu, T. Breuel, and J. Kautz, “Unsupervised image-to-image translation networks,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning . PMLR, 2017, pp. 1126–1135
2017
Earlier work this paper cites.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2223–2232
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Wulfmeier, A. Bewley, and I. Posner, “Addressing appearance change in outdoor robotics with adversarial domain adaptation,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2017, pp. 1551–1558
2017
Earlier work this paper cites.
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine, “Learning modular neural network policies for multi-task and multi-robot transfer,” in 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 2017, pp. 2169–2176
2017
Earlier work this paper cites.
D. Lopez-Paz, R. Nishihara, S. Chintala, B. Scholkopf, and L. Bottou, “Discovering causal signals in images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 6979–6987
2017
Earlier work this paper cites.
S. James, M. Bloesch, and A. J. Davison, “Task-embedded control networks for few-shot imitation learning,” in Conference on robot learning . PMLR, 2018, pp. 783–795
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto, “Robot learning in homes: Improving generalization and reducing dataset bias,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain, “Time-contrastive networks: Self-supervised learning from video,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 1134–1141
2018
Earlier work this paper cites.
K. Fang, T.-L. Wu, D. Yang, S. Savarese, and J. J. Lim, “Demo2vec: Reasoning object affordances from online videos,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 2139–2147
2018
Earlier work this paper cites.
T.-T. Do, A. Nguyen, and I. Reid, “Affordancenet: An end-to-end deep learning approach for object affordance detection,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 5882–5889
2018
Earlier work this paper cites.
G. Garcia-Hernando, S. Yuan, S. Baek, and T.-K. Kim, “First-person hand action benchmark with rgb-d videos and 3d hand pose annotations,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 409–419
2018
Earlier work this paper cites.
M. Ma, N. Marturi, Y. Li, A. Leonardis, and R. Stolkin, “Region-sequence based six-stream cnn features for general and fine-grained human action recognition in videos,” Pattern Recognition , vol. 76, pp. 506–521, 2018
2018
Earlier work this paper cites.
F. Sener and A. Yao, “Unsupervised learning and segmentation of complex activities from video,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8368–8376
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
X. B. Peng, A. Kanazawa, J. Malik, P. Abbeel, and S. Levine, “Sfv: Reinforcement learning of physical skills from videos,” ACM Transactions On Graphics (TOG) , vol. 37, no. 6, pp. 1–14, 2018
2018
Earlier work this paper cites.
Y. Liu, A. Gupta, P. Abbeel, and S. Levine, “Imitation from observation: Learning to imitate behaviors from raw video via context translation,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 1118–1125
2018
Earlier work this paper cites.
X. Huang, M.-Y. Liu, S. Belongie, and J. Kautz, “Multimodal unsupervised image-to-image translation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 172–189
2018
Earlier work this paper cites.
D. Xu, S. Nair, Y. Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese, “Neural task programming: Learning to generalize across hierarchical tasks,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 3795–3802
2018
Earlier work this paper cites.
P. Sharma, L. Mohan, L. Pinto, and A. Gupta, “Multiple interactions made easy (mime): Large scale demonstrations data for imitation,” in Conference on robot learning . PMLR, 2018, pp. 906–915
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Aytar, T. Pfaff, D. Budden, T. Paine, Z. Wang, and N. De Freitas, “Playing hard exploration games by watching youtube,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Nguyen, D. Kanoulas, L. Muratore, D. G. Caldwell, and N. G. Tsagarakis, “Translating videos to commands for robotic manipulation with deep recurrent neural networks,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 3782–3788
2018
Earlier work this paper cites.
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price et al. , “Scaling egocentric vision: The epic-kitchens dataset,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 720–736
2018
Earlier work this paper cites.
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine, “Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 3758–3765
2018
Earlier work this paper cites.
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar, “Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 3651–3657
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Caron, P. Bojanowski, J. Mairal, and A. Joulin, “Unsupervised pre-training of image features on non-curated data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 2959–2968
2019
Earlier work this paper cites.
T. Han, W. Xie, and A. Zisserman, “Video representation learning by dense predictive coding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops , 2019, pp. 0–0
2019
Earlier work this paper cites.
X. Williams and N. R. Mahapatra, “Analysis of affordance detection methods for real-world robotic manipulation,” in 2019 9th International Symposium on Embedded Computing and System Design (ISED) . IEEE, 2019, pp. 1–5
2019
Earlier work this paper cites.
B. Tekin, F. Bogo, and M. Pollefeys, “H+ o: Unified egocentric recognition of 3d hand-object poses and interactions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4511–4520
2019
Earlier work this paper cites.
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 2630–2640
2019
Earlier work this paper cites.
H. Zhang, P.-J. Lai, S. Paul, S. Kothawade, and S. Nikolaidis, “Learning collaborative action plans from youtube videos,” in The International Symposium of Robotics Research . Springer, 2019, pp. 208–223
2019
Earlier work this paper cites.
Q. Zhang, J. Chen, D. Liang, H. Liu, X. Zhou, Z. Ye, and W. Liu, “An object attribute guided framework for robot learning manipulations from human demonstration videos,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 6113–6119
2019
Cited alongside, same era.
P. Sharma, D. Pathak, and A. Gupta, “Third-person visual imitation learning via decoupled hierarchical controller,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
W. Goo and S. Niekum, “One-shot learning of multi-step tasks from observation via activity localization in auxiliary video,” in 2019 international conference on robotics and automation (ICRA) . IEEE, 2019, pp. 7755–7761
2019
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2019
Cited alongside, same era.
S. Yang, W. Zhang, W. Lu, H. Wang, and Y. Li, “Learning actions from human demonstration video for robotic manipulation,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 1805–1811
2019
Cited alongside, same era.
2019
Cited alongside, same era.
C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. Yong, J. Lee, W.-T. Chang, W. Hua, M. Georg, and M. Grundmann, “Mediapipe: A framework for perceiving and processing reality,” in Third Workshop on Computer Vision for AR/VR at IEEE Computer Vision and Pattern Recognition (CVPR) 2019 , 2019. [Online]. Available: https://mixedreality.cs.cornell.edu/s/NewTitle_May1_MediaPipe_CVPR_CV4ARVR_Workshop_2019.pdf
2019
Cited alongside, same era.
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y. Sheikh, “Openpose: Realtime multi-person 2d pose estimation using part affinity fields,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 1, pp. 172–186, 2019
2019
Cited alongside, same era.
M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg, “Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 8943–8950
2019
Cited alongside, same era.
Y. Wang, Q. Yao, J. T. Kwok, and L. M. Ni, “Generalizing from a few examples: A survey on few-shot learning,” ACM computing surveys (csur) , vol. 53, no. 3, pp. 1–34, 2020
2020
Cited alongside, same era.
S. Kadam and V. Vaidya, “Review and analysis of zero, one and few shot learning approaches,” in Intelligent Systems Design and Applications: 18th International Conference on Intelligent Systems Design and Applications (ISDA 2018) held in Vellore, India, December 6-8, 2018, Volume 1 . Springer, 2020, pp. 100–112
2020
Cited alongside, same era.
A. Méndez-Molina, E. F.Morales, and L. E. Sucar, “Causal discovery and reinforcement learning: A synergistic integration,” in Proceedings of The 11th International Conference on Probabilistic Graphical Models , ser. Proceedings of Machine Learning Research, A. Salmerón and R. Rumí, Eds., vol. 186. PMLR, 05–07 Oct 2022, pp. 421–432. [Online]. Available: https://proceedings.mlr.press/v186/mendez-molina22a.html
2022
Later among the works it cites.
A. Fishman, A. Murali, C. Eppner, B. Peele, B. Boots, and D. Fox, “Motion policy networks,” in Conference on Robot Learning . PMLR, 2023, pp. 967–977
2023
Later among the works it cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in 7th Annual Conference on Robot Learning , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu, “Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot,” in Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition@ CoRL2023 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell, “Real-world robot learning with masked visual pre-training,” in Conference on Robot Learning . PMLR, 2023, pp. 416–426
2023
Later among the works it cites.
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao, “Learning visual affordance grounding from demonstration videos,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Later among the works it cites.
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak, “Affordances from human videos as a versatile representation for robotics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 13 778–13 790
2023
Later among the works it cites.
J. Xin, L. Wang, K. Xu, C. Yang, and B. Yin, “Learning interaction regions and motion trajectories simultaneously from egocentric demonstration videos,” IEEE Robotics and Automation Letters , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang et al. , “Palm-e: An embodied multimodal language model,” 2023
2023
Later among the works it cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 11 975–11 986
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Liu, L. He, Y. Kang, Z. Zhuang, D. Wang, and H. Xu, “Ceil: Generalized contextual imitation learning,” Advances in Neural Information Processing Systems , vol. 36, pp. 75 491–75 516, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Song, T. Wang, P. Cai, S. K. Mondal, and J. P. Sahoo, “A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities,” ACM Computing Surveys , vol. 55, no. 13s, pp. 1–40, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Yang, W. Zhang, R. Song, J. Cheng, H. Wang, and Y. Li, “Watch and act: Learning robotic manipulation from visual demonstration,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , vol. 53, no. 7, pp. 4404–4416, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Xiong, H. Fu, J. Zhang, C. Bao, Q. Zhang, Y. Huang, W. Xu, A. Garg, and C. Lu, “Robotube: Learning household manipulation from human videos with simulated twin environments,” in Proceedings of The 6th Conference on Robot Learning , ser. Proceedings of Machine Learning Research, K. Liu, D. Kulic, and J. Ichnowski, Eds., vol. 205. PMLR, 14–18 Dec 2023, pp. 1–10. [Online]. Available: https://proceedings.mlr.press/v205/xiong23a.html
2023
Later among the works it cites.
Y. Li, “Deep causal learning for robotic intelligence,” Frontiers in Neurorobotics , vol. 17, p. 1128591, 2023
2023
Later among the works it cites.
T. E. Lee, S. Vats, S. Girdhar, and O. Kroemer, “Scale: Causal learning and discovery of robot manipulation skills using simulation,” in Proceedings of The 7th Conference on Robot Learning , ser. Proceedings of Machine Learning Research, J. Tan, M. Toussaint, and K. Darvish, Eds., vol. 229. PMLR, 06–09 Nov 2023, pp. 2229–2256. [Online]. Available: https://proceedings.mlr.press/v229/lee23b.html
2023
Later among the works it cites.
L. Castri, S. Mghames, and N. Bellotto, “From continual learning to causal discovery in robotics,” in Proceedings of The First AAAI Bridge Program on Continual Causality , ser. Proceedings of Machine Learning Research, M. Mundt, K. W. Cooper, D. S. Dhami, A. Ribeiro, J. S. Smith, A. Bellot, and T. Hayes, Eds., vol. 208. PMLR, 07–08 Feb 2023, pp. 85–91. [Online]. Available: https://proceedings.mlr.press/v208/castri23a.html
2023
Later among the works it cites.
2024
Closest in time.
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V. Baiyya, S. Bansal, B. Boote et al. , “Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 19 383–19 400
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. I. Erdei, T. P. Kapusi, A. Hajdu, and G. Husi, “Image-to-image translation-based deep learning application for object identification in industrial robot systems,” Robotics , vol. 13, no. 6, p. 88, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Ingelhag, J. Munkeby, J. van Haastregt, A. Varava, M. C. Welle, and D. Kragic, “A robotic skill learning system built upon diffusion policies and foundation models,” in 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN) . IEEE, 2024, pp. 748–754
2024
Closest in time.
F. Zhu, H. Wu, S. Guo, Y. Liu, C. Cheang, and T. Kong, “Mani-WM: An interactive world model for real-robot manipulation,” 2024. [Online]. Available: https://openreview.net/forum?id=aVyJwS1fqQ
2024
Closest in time.
Y. Wen, J. Lin, Y. Zhu, J. Han, H. Xu, S. Zhao, and X. Liang, “Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation,” Advances in Neural Information Processing Systems , vol. 37, pp. 41 051–41 075, 2024
2024
Closest in time.
T. E. Lee, “Causal robot learning for manipulation,” Ph.D. dissertation, Carnegie Mellon University, Pittsburgh, PA, July 2024
2024
Closest in time.
Y. M. Fourier ActionNet Team, “Actionnet: A dataset for dexterous bimanual manipulation,” 2025
2025
Closest in time.
2025
Closest in time.
H. Zhao, X. Liu, M. Xu, Y. Hao, W. Chen, and X. Han, “Taste-rob: Advancing video generation of task-oriented hand-object interaction for generalizable robotic manipulation,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 27 683–27 693
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn et al. , “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1702–1713
2025
Closest in time.
2025
Closest in time.
O. Contributors. (2010) Open Source Computer Vision Library (OpenCV). Repository commit history begins May 11, 2010 (“atomic bomb” commit):contentReference[oaicite:0]index=0; accessed 9 August 2025. [Online]. Available: https://github.com/opencv/opencv
2025
Closest in time.
2025
Closest in time.
S. Parab. (2024) The rise of diffusion models in imitation learning. Accessed 9 August 2025. [Online]. Available: https://www.trossenrobotics.com/post/the-rise-of-diffusion-models-in-imitation-learning
2025
Closest in time.
W. Yan, O. Watkins, S. James, R. Okumura, T. Darrell, and P. Abbeel. (2025) Task-specific world models for robotic manipulation. Accessed 9 August 2025. [Online]. Available: https://bcommons.berkeley.edu/task-specific-world-models-robotic-manipulation
2025
Closest in time.
2025
Closest in time.
D. Pathak, P. Mahmoudieh, G. Luo, P. Agrawal, D. Chen, Y. Shentu, E. Shelhamer, J. Malik, A. A. Efros, and T. Darrell, “Zero-shot visual imitation,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops , 2018, pp. 2050–2053
2053
Closest in time.
S. Dasari and A. Gupta, “Transformers for one-shot visual imitation,” in Conference on Robot Learning . PMLR, 2021, pp. 2071–2084
2084
Closest in time.