Fetching the paper…
Reading the bibliography…
The diffusion model is widely leveraged for either video generation or video editing.
Likert, R.: A technique for the measurement of attitudes. Archives of psychology (1932)
1932
Earlier work this paper cites.
2012
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: ICCV. pp. 4489–4497 (2015)
2015
Earlier work this paper cites.
Xu, J., Mei, T., Yao, T., Rui, Y.: Msr-vtt: A large video description dataset for bridging video and language. In: CVPR. pp. 5288–5296 (2016)
2016
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: CVPR. pp. 6299–6308 (2017)
2017
Earlier work this paper cites.
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., Gelly, S.: Fvd: A new metric for video generation. ICLR Workshop (2019)
2019
Earlier work this paper cites.
Jonathan, H., Ajay, J., Pieter, A.: Denoising diffusion probabilistic models. NeurIPS 33
2020
Earlier work this paper cites.
Saito, M., Saito, S., Koyama, M., Kobayashi, S.: Generate densely: Memory-efficient unsupervised training of high-resolution temporal gan. IJCV 128
2020
Earlier work this paper cites.
Wei, X., Zhang, T., Li, Y., Zhang, Y., Wu, F.: Multi-modality cross attention network for image and sentence matching. In: CVPR. pp. 10941–10950 (2020)
2020
Earlier work this paper cites.
Chen, C.F.R., Fan, Q., Panda, R.: Crossvit: Cross-attention multi-scale vision transformer for image classification. In: ICCV. pp. 357–366 (2021)
2021
Earlier work this paper cites.
Prafulla, D., Alexander, N.: Diffusion models beat gans on image synthesis. NeurIPS 34
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Jonathan, H., Tim, S.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
2022
Earlier work this paper cites.
Konpat, P., Nattanat, C., Suttisak, W., Supasorn, S.: Diffusion autoencoders: Toward a meaningful and decodable representation. In: CVPR. pp. 10619–10629 (2022)
2022
Earlier work this paper cites.
Nan, L., Shuang, L., Yilun, D., Antonio, T., Tenenbaum, J.B.: Compositional visual generation with composable diffusion models. In: ECCV. pp. 423–439 (2022)
2022
Earlier work this paper cites.
Parmar, G., Zhang, R., Zhu, J.Y.: On aliased resizing and surprising subtleties in gan evaluation. In: CVPR. pp. 11410–11420 (2022)
2022
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR. pp. 10684–10695 (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al.: Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2
2023
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Fidler, S.W.K.S., Kreis, K.: Align your latents: High-resolution video synthesis with latent diffusion models. In: CVPR. pp. 22563–22575 (2023)
2023
Cited alongside, same era.
Brooks, T., Holynski, A., Efros, A.A.: Instructpix2pix: Learning to follow image editing instructions. In: CVPR. pp. 18392–18402 (2023)
2023
Cited alongside, same era.
Ceylan, D., Huang, C.H.P., Mitra, N.J.: Pix2video: Video editing using image diffusion. In: ICCV. pp. 23206–23217 (2023)
2023
Cited alongside, same era.
Chai, W., Guo, X., Wang, G., Lu, Y.: Stablevideo: Text-driven consistency-aware diffusion video editing. In: ICCV. pp. 23040–23050 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Closest in time.
Ni, H., Shi, C., Li, K., Huang, S.X., Min, M.R.: Conditional image-to-video generation with latent flow diffusion models. In: CVPR. pp. 18444–18455 (2023)
2023
Closest in time.
2023
Closest in time.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In: CVPR. pp. 22500–22510 (2023)
2023
Closest in time.
Tumanyan, N., Geyer, M., Bagon, S., Dekel, T.: Plug-and-play diffusion features for text-driven image-to-image translation. In: CVPR. pp. 1921–1930 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Wu, J.Z., Ge, Y., Wang, X., Lei, W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., Shou, M.Z.: Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In: ICCV. pp. 7623–7633 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.