Fetching the paper…
Reading the bibliography…
Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control actions are subsequently derived.
H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in
2012
Earlier work this paper cites.
M. L. Puterman,
2014
Earlier work this paper cites.
Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in
2015
Earlier work this paper cites.
D. Damen, T. Leelasawassuk, and W. Mayol-Cuevas, “You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance,”
2016
Earlier work this paper cites.
S. Huang, Q. Wang, S. Zhang, S. Yan, and X. He, “Dynamic context correspondence network for semantic alignment,” in
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark,
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,”
2021
Earlier work this paper cites.
M. Janner, Q. Li, and S. Levine, “Offline reinforcement learning as one big sequence modeling problem,”
2021
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani,
2021
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Huang, L. Yang, B. He, S. Zhang, X. He, and A. Shrivastava, “Learning semantic correspondence with sparse annotations,” in
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Cited alongside, same era.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du,
2023
Later among the works it cites.
2023
Later among the works it cites.
B. He, X. Yang, H. Wang, Z. Wu, H. Chen, S. Huang, Y. Ren, S.-N. Lim, and A. Shrivastava, “Towards scalable neural representation for diverse videos,” in
2023
Later among the works it cites.
Y. Du, S. Yang, P. Florence, F. Xia, A. Wahid, brian ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. P. Kaelbling, A. Zeng, and J. Tompson, “Video language planning,” in
2024
Closest in time.
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal, “Compositional foundation models for hierarchical planning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel, “Learning universal policies via text-guided video generation,”
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
R. Zheng, C.-A. Cheng, H. D. III, F. Huang, and A. Kolobov, “PRISE: LLM-style sequence compression for learning temporal action abstractions in control,” in
2024
Closest in time.
2024
Closest in time.
S. Thakur, C. Beyan, P. Morerio, V. Murino, and A. Del Bue, “Leveraging next-active objects for context-aware anticipation in egocentric videos,” in
2024
Closest in time.
S. Huang, D.-A. Huang, Z. Yu, S. Lan, S. Radhakrishnan, J. M. Alvarez, A. Shrivastava, and A. Anandkumar, “What is point supervision worth in video instance segmentation?” in
2024
Closest in time.
S. Huang, S. Suri, K. Gupta, S. S. Rambhatla, S.-n. Lim, and A. Shrivastava, “Uvis: Unsupervised video instance segmentation,” in
2024
Closest in time.