Fetching the paper…
Reading the bibliography…
Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Layered neural atlases for consistent video editing
Kasten, Y., Ofri, D., Wang, O., and Dekel, T · 2021
Earlier work this paper cites.
Musiq: Multi-scale image quality transformer
Ke, J., Wang, Q., Wang, Y., Milanfar, P., and Yang, F · 2021
Earlier work this paper cites.
Diffedit: Diffusion-based semantic image editing with mask guidance
Couairon, G., Verbeek, J., Schwenk, H., and Cord, M · 2022
Earlier work this paper cites.
Direct inversion: Optimization-free text-driven real image editing with diffusion models
Elarabawy, A., Kamath, H., and Denton, S · 2022
Earlier work this paper cites.
Assessing a single image in reference-guided image synthesis
Guo, J., Du, C., Wang, J., Huang, H., Wan, P., and Huang, G · 2022
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Earlier work this paper cites.
aesthetic-predictor
LAION-AI · 2022
Earlier work this paper cites.
Rectified flow: A marginal preserving approach to optimal transport
Liu, Q · 2022
Earlier work this paper cites.
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Blended latent diffusion
Avrahami, O., Fried, O., and Lischinski, D · 2023
Earlier work this paper cites.
Sega: Instructing text-to-image models using semantic guidance
Brack, M., Friedrich, F., Hintersdorf, D., Struppek, L., Schramowski, P., and Kersting, K · 2023
Earlier work this paper cites.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Cao, M., Wang, X., Qi, Z., Shan, Y., Qie, X., and Zheng, Y · 2023
Earlier work this paper cites.
Stablevideo: Text-driven consistency-aware diffusion video editing
Chai, W., Guo, X., Wang, G., and Lu, Y · 2023
Earlier work this paper cites.
Control-a-video: Controllable text-to-video generation with diffusion models
Chen, W., Ji, Y., Wu, J., Wu, H., Xie, P., Li, J., Xia, X., Xiao, X., and Lin, L · 2023
Earlier work this paper cites.
Flatten: optical flow-guided attention for consistent text-to-video editing
Cong, Y., Xu, M., Simon, C., Chen, S., Ren, J., Xie, Y., Perez-Rua, J.-M., Rosenhahn, B., Xiang, T., and He, S · 2023
Earlier work this paper cites.
Diffedit: Diffusion-based semantic image editing with mask guidance
Couairon, G., Verbeek, J., Schwenk, H., and Cord, M · 2023
Cited alongside, same era.
Tokenflow: Consistent diffusion features for consistent video editing
Geyer, M., Bar-Tal, O., Bagon, S., and Dekel, T · 2023
Cited alongside, same era.
Region-aware diffusion for zero-shot text-driven image editing
Huang, N., Tang, F., Dong, W., Lee, T.-Y., and Xu, C · 2023
Cited alongside, same era.
Shape-aware text-driven layered video editing
Lee, Y.-C., Jang, J.-Z. G., Chen, Y.-T., Qiu, E., and Huang, J.-B · 2023
Cited alongside, same era.
Amt: All-pairs multi-field transforms for efficient frame interpolation
Li, Z., Zhu, Z.-L., Han, L.-H., Hou, Q., Guo, C.-L., and Cheng, M.-M · 2023
Cited alongside, same era.
On exact inversion of dpm-solvers
Hong, S., Lee, K., Jeon, S. Y., Bae, H., and Chun, S. Y · 2024
Closest in time.
Pnp inversion: Boosting diffusion-based editing with 3 lines of code
Ju, X., Zeng, A., Bian, Y., Liu, S., and Xu, Q · 2024
Closest in time.
Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models
Kara, O., Kurtkaya, B., Yesiltepe, H., Rehg, J. M., and Yanardag, P · 2024
Closest in time.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al · 2024
Closest in time.
Anyv2v: A plug-and-play framework for any video-to-video editing tasks
Ku, M., Wei, C., Ren, W., Yang, H., and Chen, W · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miyake, D., Iohara, A., Saito, Y., and Tanaka, T · 2023
Cited alongside, same era.
Null-text inversion for editing real images using guided diffusion models
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2023
Cited alongside, same era.
Zero-shot image-to-image translation
Parmar, G., Kumar Singh, K., Zhang, R., Li, Y., Lu, J., and Zhu, J.-Y · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., and Chen, Q · 2023
Cited alongside, same era.
Preditor: Text guided image editing with diffusion prior
Ravi, H., Kelkar, S., Harikumar, M., and Kale, A · 2023
Cited alongside, same era.
Plug-and-play diffusion features for text-driven image-to-image translation
Tumanyan, N., Geyer, M., Bagon, S., and Dekel, T · 2023
Cited alongside, same era.
Lee, S., Lin, Z., and Fanti, G · 2024
Closest in time.
Video-p2p: Video editing with cross-attention control
Liu, S., Zhang, Y., Li, W., Lin, Z., and Jia, J · 2024
Closest in time.
Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation
Ma, Y., Liu, H., Wang, H., Pan, H., He, Y., Yuan, J., Zeng, A., Cai, C., Shum, H.-Y., Liu, W., et al · 2024
Closest in time.
Visual instruction inversion: Image editing via image prompting
Nguyen, T., Li, Y., Ojha, U., and Lee, Y. J · 2024
Closest in time.
Codef: Content deformation fields for temporally consistent video processing
Ouyang, H., Wang, Q., Xiao, Y., Bai, Q., Zhang, J., Zheng, K., Zhou, X., Chen, Q., and Shen, Y · 2024
Closest in time.
Improving diffusion models for inverse problems using optimal posterior covariance
Peng, X., Zheng, Z., Dai, W., Xiao, N., Li, C., Zou, J., and Xiong, H · 2024
Closest in time.
Beyond first-order tweedie: Solving inverse problems using latent diffusion
Rout, L., Chen, Y., Kumar, A., Caramanis, C., Shakkottai, S., and Chu, W.-S · 2024
Closest in time.
Edit-a-video: Single video editing with object-aware consistency
Shin, C., Kim, H., Lee, C. H., Lee, S.-g., and Yoon, S · 2024
Closest in time.
Diffusion model-based video editing: A survey
Sun, W., Tu, R.-C., Liao, J., and Tao, D · 2024
Closest in time.
Hart: Efficient visual generation with hybrid autoregressive transformer
Tang, H., Wu, Y., Yang, S., Xie, E., Chen, J., Chen, J., Zhang, Z., Cai, H., Lu, Y., and Han, S · 2024
Closest in time.
Sana: Efficient high-resolution image synthesis with linear diffusion transformers
Xie, E., Chen, J., Chen, J., Cai, H., Lin, Y., Zhang, Z., Li, M., Lu, Y., and Han, S · 2024
Closest in time.
Exact diffusion inversion via bidirectional integration approximation
Zhang, G., Lewis, J. P., and Kleijn, W. B · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, March 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Closest in time.
Mvportrait: Text-guided motion and emotion control for multi-view vivid portrait animation
Lin, Y., Fung, H., Xu, J., Ren, Z., Lau, A. S., Yin, G., and Li, X · 2025
Closest in time.
Mindomni: Unleashing reasoning generation in vision language models with rgpo
Xiao, Y., Song, L., Chen, Y., Luo, Y., Chen, Y., Gan, Y., Huang, W., Li, X., Qi, X., and Shan, Y · 2025
Closest in time.