Fetching the paper…
Reading the bibliography…
We address the problem of generating long-horizon videos for robotic manipulation tasks.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Earlier work this paper cites.
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel, “Learning universal policies via text-guided video generation,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F.-Y. Wang, W. Chen, G. Song, H.-J. Ye, Y. Liu, and H. Li, “Gen-l-video: Multi-text to long video generation via temporal co-denoising,” 2023
2023
Earlier work this paper cites.
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023
2023
Earlier work this paper cites.
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel, “Any-point trajectory modeling for policy learning,” 2023
2023
Earlier work this paper cites.
S. Yin, C. Wu, H. Yang, J. Wang, X. Wang, M. Ni, Z. Yang, L. Li, S. Liu, F. Yang, J. Fu, G. Ming, L. Wang, Z. Liu, H. Li, and N. Duan, “Nuwa-xl: Diffusion over diffusion for extremely long video generation,” 2023
2023
Earlier work this paper cites.
H. Ye, J. Zhang, S. Liu, X. Han, and W. Yang, “Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models,” 2023
2023
Earlier work this paper cites.
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence, “Interactive language: Talking to robots in real time,” IEEE Robotics and Automation Letters , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Y. Du, S. Yang, P. Florence, F. Xia, A. Wahid, brian ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. P. Kaelbling, A. Zeng, and J. Tompson, “Video language planning,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu, “Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation,” 10 2024
2024
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps et al. , “Genie: Generative interactive environments,” in Forty-first International Conference on Machine Learning , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani, “Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation,” 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. Yang, Y. Du, S. K. S. Ghasemipour, J. Tompson, L. P. Kaelbling, D. Schuurmans, and P. Abbeel, “Learning interactive real-world simulators,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y. Zhou, T. Li, and Y. You, “Open-sora: Democratizing efficient video production for all,” 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
R. Henschel, L. Khachatryan, D. Hayrapetyan, H. Poghosyan, V. Tadevosyan, Z. Wang, S. Navasardyan, and H. Shi, “Streamingt2v: Consistent, dynamic, and extendable long video generation from text,” 2024
2024
Cited alongside, same era.
H. Lu, G. Yang, N. Fei, Y. Huo, Z. Lu, P. Luo, and M. Ding, “VDT: General-purpose video diffusion transformers via mask modeling,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
H. Qiu, M. Xia, Y. Zhang, Y. He, X. Wang, Y. Shan, and Z. Liu, “Freenoise: Tuning-free longer video diffusion via noise rescheduling,” in The Twelfth International Conference on Learning Representations , 2024
2024
Later among the works it cites.
Y. Lu, Y. Liang, L. Zhu, and Y. Yang, “Freelong: Training-free long video generation with spectralblend temporal attention,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
F. Ceola, L. Natale, N. Suenderhauf, and K. Rana, “LHManip: A dataset for long-horizon language-grounded manipulation tasks in cluttered tabletop environments,” in RSS 2024 Workshop: Data Generation for Robotics , 2024
2024
Later among the works it cites.
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “RDT-1b: a diffusion foundation model for bimanual manipulation,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
OpenAI, “Learning to reason with llms,” 2024, accessed: 2025-02-27. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2025
Closest in time.
DeepSeek-AI, “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025
2025
Closest in time.