Fetching the paper…
Reading the bibliography…
Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation.
Multilingual constituency parsing with self-attention and pre-training
Kitaev, N., Cao, S., and Klein, D · 2018
Earlier work this paper cites.
Compositional video synthesis with action graphs
Bar, A., Herzig, R., Wang, X., Rohrbach, A., Chechik, G., Darrell, T., and Globerson, A · 2020
Earlier work this paper cites.
Residual energy-based models for text generation
Deng, Y., Bakhtin, A., Ott, M., Szlam, A., and Ranzato, M · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Unsupervised learning of compositional energy concepts
Du, Y., Li, S., Sharma, Y., Tenenbaum, J., and Mordatch, I · 2021
Earlier work this paper cites.
Perceiver io: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., et al · 2021
Earlier work this paper cites.
Learning to compose visual relations
Liu, N., Li, S., Du, Y., Tenenbaum, J., and Torralba, A · 2021
Earlier work this paper cites.
Controllable and compositional generation with latent-space energy-based models
Nie, W., Vahdat, A., and Anandkumar, A · 2021
Earlier work this paper cites.
Is conditional generative modeling all you need for decision-making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Instruction-driven history-aware policies for robotic manipulations
Guhur, P.-L., Chen, S., Pinel, R. G., Tapaswi, M., Laptev, I., and Schmid, C · 2023
Later among the works it cites.
Diffusion-based generation, optimization, and planning in 3d scenes
Huang, S., Wang, Z., Li, P., Jia, B., Liu, T., Zhu, Y., Liang, W., and Zhu, S.-C · 2023
Later among the works it cites.
Learning to act from actionless videos through dense correspondences, 2023
Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B · 2023
Later among the works it cites.
Adaptdiffuser: Diffusion models as adaptive self-evolving planners
Liang, Z., Mu, Y., Ding, M., Ni, F., Tomizuka, M., and Luo, P · 2023
Later among the works it cites.
Composable text controls in latent space with ODEs
Liu, G., Feng, Z., Gao, Y., Yang, Z., Liang, X., Bao, J., He, X., Cui, S., Li, Z., and Hu, Z · 2023
Later among the works it cites.
Structdiffusion: Language-guided creation of physically-valid structures using unseen objects
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modular action concept grounding in semantic video prediction
Yu, W., Chen, W., Yin, S., Easterbrook, S., and Garg, A · 2022
Cited alongside, same era.
Lad: Language augmented diffusion for reinforcement learning
Zhang, E., Lu, Y., Wang, W. Y., and Zhang, A · 2022
Cited alongside, same era.
Compositional foundation models for hierarchical planning
Ajay, A., Han, S., Du, Y., Li, S., Gupta, A., Jaakkola, T., Tenenbaum, J., Kaelbling, L., Srivastava, A., and Agrawal, P · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Cited alongside, same era.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Cited alongside, same era.
Energy-based models as zero-shot planners for compositional scene rearrangement
Gkanatsios, N., Jain, A., Xian, Z., Zhang, Y., Atkeson, C., and Fragkiadaki, K · 2023
Cited alongside, same era.
Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and mcmc
Du, Y., Durkan, C., Strudel, R., Tenenbaum, J. B., Dieleman, S., Fergus, R., Sohl-Dickstein, J., Doucet, A., and Grathwohl, W. S
Cited in the paper.
Liu, W., Du, Y., Hermans, T., Chernova, S., and Paxton, C · 2023
Later among the works it cites.
Imitating human behaviour with diffusion models
Pearce, T., Rashid, T., Kanervisto, A., Bignell, D., Sun, M., Georgescu, R., Macua, S. V., Tan, S. Z., Momennejad, I., Hofmann, K., et al · 2023
Later among the works it cites.
Exploring compositional visual generation with latent classifier guidance
Shi, C., Ni, H., Li, K., Han, S., Liang, M., and Min, M. R · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Adaptive online replanning with diffusion models
Zhou, S., Du, Y., Zhang, S., Xu, M., Shen, Y., Xiao, W., Yeung, D.-Y., and Gan, C · 2023
Later among the works it cites.
Instruct-imagen: Image generation with multi-modal instruction
Hu, H., Chan, K. C., Su, Y.-C., Chen, W., Li, Y., Sohn, K., Zhao, Y., Ben, X., Gong, B., Cohen, W., et al · 2024
Closest in time.