Fetching the paper…
Reading the bibliography…
We introduce Videoshop, a training-free video editing algorithm for localized semantic edits.
Mackay, W., Pagani, D.: Video mosaic: Laying out time in a physical space. In: Proceedings of the second ACM international conference on Multimedia. pp. 165–172 (1994)
1994
Earlier work this paper cites.
Gutmann, M., Hyvärinen, A.: Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics. pp. 297–304. JMLR Workshop and Conference Proceedings (2010)
2010
Earlier work this paper cites.
Vincent, P.: A connection between score matching and denoising autoencoders. Neural computation 23
2011
Earlier work this paper cites.
Santosa, S., Chevalier, F., Balakrishnan, R., Singh, K.: Direct space-time trajectory control for visual media editing. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. pp. 1149–1158 (2013)
2013
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. Advances in neural information processing systems 27
2014
Earlier work this paper cites.
Chan, C., Ginosar, S., Zhou, T., Efros, A.A.: Everybody dance now (2019)
2019
Earlier work this paper cites.
Karras, T., Laine, S., Aila, T.: A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4401–4410 (2019)
2019
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models (2020)
2020
Earlier work this paper cites.
Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow (2020)
2020
Earlier work this paper cites.
Kasten, Y., Ofri, D., Wang, O., Dekel, T.: Layered neural atlases for consistent video editing (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision (2021)
2021
Earlier work this paper cites.
Bhattad, A., Forsyth, D.A.: Cut-and-paste object insertion by enabling deep image prior for reshading. In: 2022 International Conference on 3D Vision (3DV). pp. 332–341. IEEE (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Karras, T., Aittala, M., Aila, T., Laine, S.: Elucidating the design space of diffusion-based generative models (2022)
2022
Earlier work this paper cites.
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., Cohen-Or, D.: Null-text inversion for editing real images using guided diffusion models (2022)
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models (2022)
2022
Earlier work this paper cites.
Wallace, B., Gokul, A., Naik, N.: Edict: Exact diffusion inversion via coupled transformations (2022)
2022
Earlier work this paper cites.
Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., Guo, B.: Advancing high-resolution video-language representation with large-scale video transcriptions (2022)
2022
Earlier work this paper cites.
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets (2023)
2023
Earlier work this paper cites.
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S.W., Fidler, S., Kreis, K.: Align your latents: High-resolution video synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 22563–22575 (2023)
2023
Cited alongside, same era.
Ceylan, D., Huang, C.H.P., Mitra, N.J.: Pix2video: Video editing using image diffusion (2023)
2023
Cited alongside, same era.
Chang, S.Y., Chen, H.T., Liu, T.L.: Diffusionatlas: High-fidelity consistent diffusion video editing (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Couairon, P., Rambour, C., Haugeard, J.E., Thome, N.: Videdit: Zero-shot and spatially aware text-driven video editing (2023)
Yang, S., Zhou, Y., Liu, Z., Loy, C.C.: Rerender a video: Zero-shot text-guided video-to-video translation (2023)
2023
Later among the works it cites.
Yatim, D., Fridman, R., Bar-Tal, O., Kasten, Y., Dekel, T.: Space-time diffusion features for zero-shot text-driven motion transfer (2023)
2023
Later among the works it cites.
Yin, W., Yin, H., Baraka, K., Kragic, D., Björkman, M.: Dance style transfer with cross-modal transformer (2023)
2023
Later among the works it cites.
Zhang, D.J., Wu, J.Z., Liu, J.W., Zhao, R., Ran, L., Gu, Y., Gao, D., Shou, M.Z.: Show-1: Marrying pixel and latent diffusion models for text-to-video generation (2023)
2023
Later among the works it cites.
Zhang, G., Lewis, J.P., Kleijn, W.B.: Exact diffusion inversion via bi-directional integration approximation (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Decatur, D., Lang, I., Hanocka, R.: 3d highlighter: Localizing regions on 3d shapes via text descriptions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20930–20939 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Geyer, M., Bar-Tal, O., Bagon, S., Dekel, T.: Tokenflow: Consistent diffusion features for consistent video editing (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Hu, Y., Liu, B., Kasai, J., Wang, Y., Ostendorf, M., Krishna, R., Smith, N.A.: Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering (2023)
2023
Cited alongside, same era.
Kang, M., Zhu, J.Y., Zhang, R., Park, J., Shechtman, E., Paris, S., Park, T.: Scaling up gans for text-to-image synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10124–10134 (2023)
2023
Cited alongside, same era.
Zhang, K., Mo, L., Chen, W., Sun, H., Su, Y.: Magicbrush: A manually annotated dataset for instruction-guided image editing (2023)
2023
Later among the works it cites.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3836–3847 (2023)
2023
Later among the works it cites.
Zuo, Z., Zhang, Z., Luo, Y., Zhao, Y., Zhang, H., Yang, Y., Wang, M.: Cut-and-paste: Subject-driven video editing with attention control (2023)
2023
Later among the works it cites.
Bhattad, A., McKee, D., Hoiem, D., Forsyth, D.: Stylegan knows normal, depth, albedo, and more. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Bhattad, A., Soole, J., Forsyth, D.: Stylitgan: Image-based relighting via latent control. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4231–4240 (2024)
2024
Closest in time.
Chen, H., Zhang, Y., Cun, X., Xia, M., Wang, X., Weng, C., Shan, Y.: Videocrafter2: Overcoming data limitations for high-quality video diffusion models (2024)
2024
Closest in time.
Inc., A.: Final cut pro. https://www.apple.com/final-cut-pro/ (2023), accessed: 2024-03-03
2024
Closest in time.
Incorporated, A.S.: Adobe premiere pro. https://www.adobe.com/products/premiere.html (2023), accessed: 2024-03-03
2024
Closest in time.
Jeong, H., Ye, J.C.: Ground-a-video: Zero-shot grounded video editing using text-to-image diffusion models (2024)
2024
Closest in time.
Kahatapitiya, K., Karjauv, A., Abati, D., Porikli, F., Asano, Y.M., Habibian, A.: Object-centric diffusion for efficient video editing (2024)
2024
Closest in time.
2024
Closest in time.
Michel, O., Bhattad, A., VanderBilt, E., Krishna, R., Kembhavi, A., Gupta, T.: Object 3dit: Language-guided 3d-aware image editing. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Ren, Y., Zhou, Y., Yang, J., Shi, J., Liu, D., Liu, F., Kwon, M., Shrivastava, A.: Customize-a-video: One-shot motion customization of text-to-video diffusion models (2024)
2024
Closest in time.
Sarkar, A., Mai, H., Mahapatra, A., Lazebnik, S., Forsyth, D.A., Bhattad, A.: Shadows don’t lie and lines can’t bend! generative models don’t know projective geometry… for now. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 28140–28149 (2024)
2024
Closest in time.
Wang, F.Y., Huang, Z., Shi, X., Bian, W., Song, G., Liu, Y., Li, H.: Animatelcm: Accelerating the animation of personalized diffusion models and adapters with decoupled consistency learning (2024)
2024
Closest in time.