Fetching the paper…
Reading the bibliography…
Diffusion-based video editing have reached impressive quality and can transform either the global style, local structure, and attributes of given video inputs, following textual edit prompts.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial networks. Neural Information Processing Systems (2014)
2014
Earlier work this paper cites.
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. International Conference on Learning Representations (2014)
2014
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) (2015)
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Neural Information Processing Systems (2020)
2020
Earlier work this paper cites.
Qin, X., Zhang, Z., Huang, C., Dehghan, M., Zaiane, O.R., Jagersand, M.: U2-net: Going deeper with nested u-structure for salient object detection. Pattern Recognition (2020)
2020
Earlier work this paper cites.
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., Koltun, V.: Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020)
2020
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (2021)
2021
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. International Conference on Learning Representations (2021)
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Ho, J., Saharia, C., Chan, W., Fleet, D.J., Norouzi, M., Salimans, T.: Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research (2022)
2022
Earlier work this paper cites.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., Fleet, D.J.: Video diffusion models. Neural Information Processing Systems (2022)
2022
Earlier work this paper cites.
Li, M., Lin, J., Meng, C., Ermon, S., Han, S., Zhu, J.Y.: Efficient spatially sparse inference for conditional gans and diffusion models. Neural Information Processing Systems (2022)
2022
Earlier work this paper cites.
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., Zhu, J.: Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Neural Information Processing Systems (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.Y., Ermon, S.: Sdedit: Image synthesis and editing with stochastic differential equations. International Conference on Learning Representations (2022)
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2022)
2022
Cited alongside, same era.
Salimans, T., Ho, J.: Progressive distillation for fast sampling of diffusion models. International Conference on Learning Representations (2022)
2022
Cited alongside, same era.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al.: Laion-5b: An open large-scale dataset for training next generation image-text models. Neural Information Processing Systems (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Meng, C., Rombach, R., Gao, R., Kingma, D., Ermon, S., Ho, J., Salimans, T.: On distillation of guided diffusion models. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2023)
2023
Later among the works it cites.
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., Cohen-Or, D.: Null-text inversion for editing real images using guided diffusion models. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2023)
2023
Later among the works it cites.
Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: IEEE International Conference on Computer Vision (2023)
2023
Later among the works it cites.
Qi, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., Chen, Q.: Fatezero: Fusing attentions for zero-shot text-based video editing. IEEE International Conference on Computer Vision (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Avrahami, O., Fried, O., Lischinski, D.: Blended latent diffusion. ACM Transactions on Graphics (TOG) (2023)
2023
Cited alongside, same era.
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S.W., Fidler, S., Kreis, K.: Align your latents: High-resolution video synthesis with latent diffusion models. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2023)
2023
Cited alongside, same era.
Bolya, D., Fu, C.Y., Dai, X., Zhang, P., Feichtenhofer, C., Hoffman, J.: Token merging: Your vit but faster. International Conference on Learning Representations (2023)
2023
Cited alongside, same era.
Bolya, D., Hoffman, J.: Token merging for fast stable diffusion. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Esser, P., Chiu, J., Atighehchian, P., Granskog, J., Germanidis, A.: Structure and content-guided video synthesis with diffusion models. In: IEEE International Conference on Computer Vision (2023)
2023
Cited alongside, same era.
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., Cohen-Or, D.: Prompt-to-prompt image editing with cross attention control. International Conference on Learning Representations (2023)
2023
Cited alongside, same era.
Hong, W., Ding, M., Zheng, W., Liu, X., Tang, J.: Cogvideo: Large-scale pretraining for text-to-video generation via transformers. International Conference on Learning Representations (2023)
2023
Cited alongside, same era.
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al.: Make-a-video: Text-to-video generation without text-video data. International Conference on Learning Representations (2023)
2023
Later among the works it cites.
Tumanyan, N., Geyer, M., Bagon, S., Dekel, T.: Plug-and-play diffusion features for text-driven image-to-image translation. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2023)
2023
Later among the works it cites.
Wu, J.Z., Ge, Y., Wang, X., Lei, S.W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., Shou, M.Z.: Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In: IEEE International Conference on Computer Vision (2023)
2023
Later among the works it cites.
Yun, Y.K., Lin, W.: Selfreformer: Self-refined network with transformer for salient object detection. IEEE Transactions on Multimedia (2023)
2023
Later among the works it cites.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: IEEE International Conference on Computer Vision (2023)
2023
Later among the works it cites.
Geyer, M., Bar-Tal, O., Bagon, S., Dekel, T.: Tokenflow: Consistent diffusion features for consistent video editing. International Conference on Learning Representations (2024)
2024
Closest in time.
Kim, D., Lai, C.H., Liao, W.H., Murata, N., Takida, Y., Uesaka, T., He, Y., Mitsufuji, Y., Ermon, S.: Consistency trajectory models: Learning probability flow ode trajectory of diffusion. International Conference on Learning Representations (2024)
2024
Closest in time.
Li, X., Ma, C., Yang, X., Yang, M.H.: Vidtome: Video token merging for zero-shot video editing. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2024)
2024
Closest in time.
Liu, S., Zhang, Y., Li, W., Lin, Z., Jia, J.: Video-p2p: Video editing with cross-attention control. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (2024)
2024
Closest in time.
2024
Closest in time.
Zhang, Y., Wei, Y., Jiang, D., Zhang, X., Zuo, W., Tian, Q.: Controlvideo: Training-free controllable text-to-video generation. International Conference on Learning Representations (2024)
2024
Closest in time.