Fetching the paper…
Reading the bibliography…
Diffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control.
M. Welling and Y. W. Teh, “Bayesian learning via stochastic gradient langevin dynamics,” in Proceedings of the 28th international conference on machine learning (ICML-11) . Citeseer, 2011, pp. 681–688
2011
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4401–4410
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on robot learning . PMLR, 2020, pp. 1094–1100
2020
Earlier work this paper cites.
Y. Zhou, B. Karimi, J. Yu, Z. Xu, and P. Li, “Towards better generalization of adaptive gradient methods,” Advances in Neural Information Processing Systems , vol. 33, pp. 810–821, 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy
2021
Earlier work this paper cites.
A. Akbari, M. Awais, M. Bashar, and J. Kittler, “How does loss function affect generalization performance of deep learning? application to human age estimation,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 141–151. [Online]. Available: https://proceedings.mlr.press/v139/akbari21a.html
2021
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in neural information processing systems , vol. 35, pp. 36 479–36 494, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 12 104–12 113
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Liu, T. Hermans, S. Chernova, and C. Paxton, “Structdiffusion: Object-centric diffusion for semantic rearrangement of novel objects,” in Workshop on Language and Robotics at CoRL 2022 , 2022
2022
Earlier work this paper cites.
Y. Zhao, H. Zhang, and X. Hu, “Penalizing gradient norm for efficiently improving generalization in deep learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 26 982–26 992
2022
Earlier work this paper cites.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Earlier work this paper cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” RSS , 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin et al. , “Scaling vision transformers to 22 billion parameters,” in International Conference on Machine Learning . PMLR, 2023, pp. 7480–7512
2023
Cited alongside, same era.
H. Ha, P. Florence, and S. Song, “Scaling up and distilling down: Language-guided robot skill acquisition,” in Conference on Robot Learning . PMLR, 2023, pp. 3766–3777
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Seo, D. Hafner, H. Liu, F. Liu, S. James, K. Lee, and P. Abbeel, “Masked world models for visual control,” in Conference on Robot Learning . PMLR, 2023, pp. 1332–1344
2023
Later among the works it cites.
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh, “Video generation models as world simulators,” 2024. [Online]. Available: https://openai.com/research/video-generation-models-as-world-simulators
2024
Closest in time.
K. Lee, K. Sohn, and J. Shin, “Dreamflow: High-quality text-to-3d generation by approximating probability flow,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=GURqUuTebY
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang, “Visual reinforcement learning with self-supervised 3d representations,” IEEE Robotics and Automation Letters , vol. 8, no. 5, pp. 2890–2897, 2023
2023
Cited alongside, same era.
Z. Xian, N. Gkanatsios, T. Gervet, and K. Fragkiadaki, “Unifying diffusion models with action detection transformers for multi-task robotic manipulation,” in 7th Annual Conference on Robot Learning , 2023
2023
Cited alongside, same era.
Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang, “Gnfactor: Multi-task real robot learning with generalizable neural feature fields,” in Conference on Robot Learning . PMLR, 2023, pp. 284–301
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang, “Visual reinforcement learning with self-supervised 3d representations,” IEEE Robotics and Automation Letters , vol. 8, no. 5, pp. 2890–2897, 2023
2023
Cited alongside, same era.
Closest in time.
2024
Closest in time.
J. Brehmer, J. Bose, P. De Haan, and T. S. Cohen, “Edgi: Equivariant diffusion for planning with embodied agents,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
K. Lee, S. Kim, and J. Choi, “Refining diffusion planner for reliable behavior synthesis by automatic detection of infeasible plans,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
S. Zhou, Y. Du, S. Zhang, M. Xu, Y. Shen, W. Xiao, D.-Y. Yeung, and C. Gan, “Adaptive online replanning with diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov, “Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,” in First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024
2024
Closest in time.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,” in ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation , 2024
2024
Closest in time.
2024
Closest in time.
Y. Zhu, Z. Ou, X. Mou, and J. Tang, “Retrieval-augmented embodied agents,” 2024
2024
Closest in time.
S. Yang, Y. Ze, and H. Xu, “Movie: Visual model-based policy adaptation for view generalization,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Ze, Y. Liu, R. Shi, J. Qin, Z. Yuan, J. Wang, and H. Xu, “H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
T. Wu, M. Wu, J. Zhang, Y. Gan, and H. Dong, “Learning score-based grasping primitive for human-assisting dexterous grasping,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Wen, Y. Zhu, M. Zhu, J. Li, Z. Xu et al. , “Object-centric instruction augmentation for robotic manipulation,” 2024
2024
Closest in time.
M. Zhu, Y. Zhu, J. Li, J. Wen, Z. Xu et al. , “Language-conditioned robotic manipulation with fast and slow thinking,” 2024
2024
Closest in time.