Fetching the paper…
Reading the bibliography…
Robot imitation learning relies on 4D multi-view sequential images.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PmLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Brooks, J. Hellsten, M. Aittala, T.-C. Wang, T. Aila, J. Lehtinen, M.-Y. Liu, A. Efros, and T. Karras, “Generating long videos of dynamic scenes,” Advances in Neural Information Processing Systems , vol. 35, pp. 31 769–31 781, 2022
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” Advances in Neural Information Processing Systems , vol. 35, pp. 8633–8646, 2022
2022
Earlier work this paper cites.
T. Yu, T. Xiao, J. Tompson, A. Stone, S. Wang, A. Brohan, J. Singh, C. Tan, D. M, J. Peralta, K. Hausman, B. Ichter, and F. Xia, “Scaling Robot Learning with Semantically Imagined Experience,” in Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Zeng, H. Zhang, L. Gao, X. Li, J. Qian, and H. T. Shen, “Visual commonsense-aware representation network for video captioning,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Earlier work this paper cites.
I. Kapelyukh, V. Vosylius, and E. Johns, “Dall-e-bot: Introducing web-scale diffusion models to robotics,” IEEE Robotics and Automation Letters , vol. 8, no. 7, pp. 3956–3963, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 4015–4026
2023
Earlier work this paper cites.
R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick, “Zero-1-to-3: Zero-shot one image to 3d object,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 9298–9309
2023
Earlier work this paper cites.
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
R. Gao, K. Chen, E. Xie, L. Hong, Z. Li, D.-Y. Yeung, and Q. Xu, “MagicDrive: Street view generation with diverse 3d geometry control,” in International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
A. Swerdlow, R. Xu, and B. Zhou, “Street-view image generation from a bird’s-eye view layout,” IEEE Robotics Autom. Lett. , vol. 9, no. 4, pp. 3578–3585, 2024
2024
Cited alongside, same era.
Y. Deng, R. Wang, Y. Zhang, Y.-W. Tai, and C.-K. Tang, “Dragvideo: Interactive drag-style video editing,” in European Conference on Computer Vision . Springer, 2024, pp. 183–199
2024
Cited alongside, same era.
S. Liu, Y. Zhang, W. Li, Z. Lin, and J. Jia, “Video-p2p: Video editing with cross-attention control,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 8599–8608
2024
Cited alongside, same era.
N. Sharma, A. Tripathi, A. Chakraborty, and A. Mishra, “Sketch-guided image inpainting with partial discrete diffusion process,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 6024–6034
2024
Later among the works it cites.
X. He, F. Yang, F. Liu, and G. Lin, “Few-shot image generation via style adaptation and content preservation,” IEEE Transactions on Neural Networks and Learning Systems , 2024
2024
Later among the works it cites.
Z. Huang, H. Wen, J. Dong, Y. Wang, Y. Li, X. Chen, Y.-P. Cao, D. Liang, Y. Qiao, B. Dai et al. , “Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9784–9794
2024
Later among the works it cites.
H. Lin, Y. Chen, J. Wang, W. An, M. Wang, F. Tian, Y. Liu, G. Dai, J. Wang, and Q. Wang, “Schedule your edit: A simple yet effective diffusion noise schedule for image editing,” Advances in Neural Information Processing Systems , vol. 37, pp. 115 712–115 756, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Yi, T. Hu, M. Xia, Y. Tang, and Y.-J. Liu, “Feditnet++: Few-shot editing of latent semantics in gan spaces with correlated attribute disentanglement,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Cited alongside, same era.
J. Wu, J.-W. Bian, X. Li, G. Wang, I. Reid, P. Torr, and V. A. Prisacariu, “Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting editing,” in European Conference on Computer Vision . Springer, 2024, pp. 55–71
2024
Cited alongside, same era.
Y. Zhou, D. Zhou, M.-M. Cheng, J. Feng, and Q. Hou, “Storydiffusion: Consistent self-attention for long-range image and video generation,” Advances in Neural Information Processing Systems , vol. 37, pp. 110 315–110 340, 2024
2024
Cited alongside, same era.
Y. Liu, X. Cun, X. Liu, X. Wang, Y. Zhang, H. Chen, Y. Liu, T. Zeng, R. Chan, and Y. Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 139–22 149
2024
Cited alongside, same era.
Z. You, Z. Wen, Y. Chen, X. Li, R. Zeng, Y. Wang, and M. Tan, “Towards long video understanding via fine-detailed video story generation,” IEEE Transactions on Circuits and Systems for Video Technology , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Chen, Z. Mandi, H. Bharadhwaj, M. Sharma, S. Song, A. Gupta, and V. Kumar, “Semantically controllable augmentations for generalizable robot learning,” The International Journal of Robotics Research , p. 02783649241273686, 2024
2024
Cited alongside, same era.
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar, “Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 4788–4795
2024
Cited alongside, same era.
2024
Later among the works it cites.
Y. Huang, J. Huang, Y. Liu, M. Yan, J. Lv, J. Liu, W. Xiong, H. Zhang, L. Cao, and S. Chen, “Diffusion model-based image editing: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2025
2025
Closest in time.
J. Kwon, H. Cho, and J. Kim, “Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 26 855–26 865
2025
Closest in time.
J. Ni, Y. Guo, Y. Liu, R. Chen, L. Lu, and Z. Wu, “Maskgwm: A generalizable driving world model with video mask reconstruction,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 22 381–22 391
2025
Closest in time.
2025
Closest in time.
Z. Shi, D. Huo, Y. Zhou, Y. Min, J. Lu, and X. Zuo, “Imfine: 3d inpainting via geometry-guided multi-view refinement,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 26 694–26 703
2025
Closest in time.
2025
Closest in time.
M. Sun, W. Wang, G. Li, J. Liu, J. Sun, W. Feng, S. Lao, S. Zhou, Q. He, and J. Liu, “Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 7364–7373
2025
Closest in time.
D. Xie, Z. Xu, Y. Hong, H. Tan, D. Liu, F. Liu, A. Kaufman, and Y. Zhou, “Progressive autoregressive video diffusion models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 6322–6332
2025
Closest in time.
T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang, “From slow bidirectional to fast autoregressive video diffusion models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 22 963–22 974
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Mu, T. Chen, S. Peng, Z. Chen, Z. Gao, Y. Zou, L. Lin, Z. Xie, and P. Luo, “Robotwin: Dual-arm robot benchmark with generative digital twins (early version),” in European Conference on Computer Vision . Springer, 2025, pp. 264–273
2025
Closest in time.