Fetching the paper…
Reading the bibliography…
In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation.
Learning a confidence measure for optical flow
Mac Aodha, O., Humayun, A., Pollefeys, M., and Brostow, G. J · 2012
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Koizumi, Y., Ohishi, Y., Niizumi, D., Takeuchi, D., and Yasuda, M · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Teed, Z. and Deng, J · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss
Cheng, X., Lin, H., Wu, X., Yang, F., and Shen, D · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Earlier work this paper cites.
Learning accurate dense correspondences and when to trust them
Truong, P., Danelljan, M., Van Gool, L., and Timofte, R · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ding, M., Zheng, W., Hong, W., and Tang, J · 2022
Cited alongside, same era.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Later among the works it cites.
Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation
Wang, H., Du, X., Li, J., Yeh, R. A., and Shakhnarovich, G · 2022
Later among the works it cites.
Nüwa: Visual synthesis pre-training for neural visual world creation
Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N · 2022
Later among the works it cites.
Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models
Xu, J., Wang, X., Cheng, W., Cao, Y.-P., Shan, Y., Qie, X., and Gao, S · 2022
Later among the works it cites.
Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Magic3d: High-resolution text-to-3d content creation
Lin, C.-H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.-Y., and Lin, T.-Y · 2022
Cited alongside, same era.
Pseudo numerical methods for diffusion models on manifolds
Liu, L., Ren, Y., Lin, Z., and Zhao, Z · 2022
Cited alongside, same era.
Latent-nerf for shape-guided generation of 3d shapes and textures
Metzer, G., Richardson, E., Patashnik, O., Giryes, R., and Cohen-Or, D · 2022
Cited alongside, same era.
Introducing chatgpt, 2022
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
An, J., Zhang, S., Yang, H., Gupta, S., Huang, J.-B., Luo, J., and Yin, X · 2023
Closest in time.
Align your latents: High-resolution video synthesis with latent diffusion models
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K · 2023
Closest in time.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2023
Closest in time.
Structure and content-guided video synthesis with diffusion models
Esser, P., Chiu, J., Atighehchian, P., Granskog, J., and Germanidis, A · 2023
Closest in time.
Palm 2 technical report, 2023
Google · 2023
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Closest in time.
Debiasing scores and prompts of 2d diffusion for robust text-to-3d generation
Hong, S., Ahn, D., and Kim, S · 2023
Closest in time.
Free-bloom: Zero-shot text-to-video generator with llm director and ldm animator
Huang, H., Feng, Y., Shi, C., Xu, L., Yu, J., and Yang, S · 2023
Closest in time.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., and Shi, H · 2023
Closest in time.
Videofusion: Decomposed diffusion models for high-quality video generation
Luo, Z., Chen, D., Zhang, Y., Huang, Y., Wang, L., Shen, Y., Zhao, D., Zhou, J., and Tan, T · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., and Chen, Q · 2023
Closest in time.
Large language models are human-level prompt engineers
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J · 2023
Closest in time.