Fetching the paper…
Reading the bibliography…
We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthesis.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
Shot durations, shot classes, and the increased pace of popular movies, 2015
Cutting, J. E. and Candan, A · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
3d u-net: learning dense volumetric segmentation from sparse annotation
Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., and Ronneberger, O · 2016
Earlier work this paper cites.
Improved techniques for training GANs
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Style transfer by relaxed optimal transport and self-similarity
Kolkin, N., Salavon, J., and Shakhnarovich, G · 2019
Earlier work this paper cites.
Effectively unbiased FID and Inception Score and where to find them
Chong, M. J. and Forsyth, D · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Resolution dependent GAN interpolation for controllable image synthesis between domains
Pinkney, J. N. and Adler, D · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal GAN
Saito, M., Saito, S., Koyama, M., and Kobayashi, S · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
CogVideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Cited alongside, same era.
Breathing life into sketches using text-to-video priors
Gal, R., Vinker, Y., Alaluf, Y., Bermano, A. H., Cohen-Or, D., Shamir, A., and Chechik, G · 2023
Later among the works it cites.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.-B., Liu, M.-Y., and Balaji, Y · 2023
Later among the works it cites.
Emu Video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Later among the works it cites.
Gu, J., Zhai, S., Zhang, Y., Susskind, J., and Jaitly, N · 2023
Later among the works it cites.
AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SDEdit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2022
Cited alongside, same era.
On aliased resizing and surprising subtleties in gan evaluation
Parmar, G., Zhang, R., and Zhu, J.-Y · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Palette: Image-to-image diffusion models
Saharia, C., Chan, W., Chang, H., Lee, C., Ho, J., Salimans, T., Fleet, D., and Norouzi, M · 2022
Cited alongside, same era.
Make-a-Video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
Nüwa: Visual synthesis pre-training for neural visual world creation
Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N · 2022
Cited alongside, same era.
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Later among the works it cites.
Photorealistic video generation with diffusion models
Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Fei-Fei, L., Essa, I., Jiang, L., and Lezama, J · 2023
Later among the works it cites.
Simple diffusion: End-to-end diffusion for high resolution images
Hoogeboom, E., Heek, J., and Salimans, T · 2023
Later among the works it cites.
VideoPoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Later among the works it cites.
https://www.pika.art/ , 2023
Pika labs · 2023
Later among the works it cites.
State of the art on diffusion models for visual computing
Po, R., Yifan, W., Golyanik, V., Aberman, K., Barron, J. T., Bermano, A. H., Chan, E. R., Dekel, T., Holynski, A., Kanazawa, A., et al · 2023
Later among the works it cites.
DreamFusion: Text-to-3D using 2D diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2023
Later among the works it cites.
StyleDrop: Text-to-image generation in any style
Sohn, K., Ruiz, N., Lee, K., Chin, D. C., Blok, I., Chang, H., Barber, J., Jiang, L., Entis, G., Li, Y., et al · 2023
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2023
Later among the works it cites.
Inflation with diffusion: Efficient temporal adaptation for text-to-video super-resolution, 2024
Yuan, X., Baek, J., Xu, K., Tov, O., and Fei, H · 2024
Closest in time.