Fetching the paper…
Reading the bibliography…
The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality.
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, “Plug-and-play diffusion features for text-driven image-to-image translation,” in CVPR . IEEE, 2023, pp. 1921–1930
1930
Earlier work this paper cites.
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J. Zhu, “Multi-concept customization of text-to-image diffusion,” in CVPR . IEEE, 2023, pp. 1931–1941
1941
Earlier work this paper cites.
B. K. P. Horn and B. G. Schunck, “Determining optical flow,” Artif. Intell. , vol. 17, no. 1-3, pp. 185–203, 1981
1981
Earlier work this paper cites.
A. Hyvärinen, “Estimation of non-normalized statistical models by score matching,” JMLR , vol. 6, pp. 695–709, 2005
2005
Earlier work this paper cites.
P. Sand and S. J. Teller, “Particle video: Long-range motion estimation using point trajectories,” in CVPR (2) . IEEE Computer Society, 2006, pp. 2195–2202
2006
Earlier work this paper cites.
D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” in ECCV , vol. 7577. Springer, 2012, pp. 611–625
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in ICLR , 2014
2014
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in ECCV , vol. 8693. Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML , vol. 37. JMLR.org, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
D. J. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in ICML , vol. 37. JMLR.org, 2015, pp. 1530–1538
2015
Earlier work this paper cites.
A. Dosovitskiy, P. Fischer, E. Ilg, P. Häusser, C. Hazirbas, V. Golkov, P. van der Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” in ICCV . IEEE Computer Society, 2015, pp. 2758–2766
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Xie and Z. Tu, “Holistically-nested edge detection,” in ICCV . IEEE Computer Society, 2015, pp. 1395–1403
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Generating videos with scene dynamics,” in NIPS , 2016, pp. 613–621
2016
Earlier work this paper cites.
N. Mayer, E. Ilg, P. Häusser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” in CVPR . IEEE Computer Society, 2016, pp. 4040–4048
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” CoRR , vol. abs/1606.08415, 2016
2016
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in NIPS , 2017, pp. 6306–6315
2017
Earlier work this paper cites.
X. Huang and S. J. Belongie, “Arbitrary style transfer in real-time with adaptive instance normalization,” in ICCV . IEEE Computer Society, 2017, pp. 1510–1519
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Saito, E. Matsumoto, and S. Saito, “Temporal generative adversarial nets with singular value clipping,” in ICCV . IEEE Computer Society, 2017, pp. 2849–2858
2017
Earlier work this paper cites.
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville, “Film: Visual reasoning with a general conditioning layer,” in AAAI . AAAI Press, 2018, pp. 3942–3951
2018
Earlier work this paper cites.
Y. Wu and K. He, “Group normalization,” in ECCV (13) , vol. 11217. Springer, 2018, pp. 3–19
2018
Earlier work this paper cites.
S. Tulyakov, M. Liu, X. Yang, and J. Kautz, “Mocogan: Decomposing motion and content for video generation,” in CVPR . IEEE, 2018, pp. 1526–1535
2018
Earlier work this paper cites.
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, “Spectral normalization for generative adversarial networks,” in ICLR , 2018
2018
Earlier work this paper cites.
W. Lai, J. Huang, O. Wang, E. Shechtman, E. Yumer, and M. Yang, “Learning blind video temporal consistency,” in ECCV , vol. 11219. Springer, 2018, pp. 179–195
2018
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in NeurIPS , 2019, pp. 11 895–11 907
2019
Earlier work this paper cites.
J. Lin, C. Gan, and S. Han, “TSM: temporal shift module for efficient video understanding,” in ICCV . IEEE, 2019, pp. 7082–7092
2019
Earlier work this paper cites.
Y. Ge, R. Zhang, X. Wang, X. Tang, and P. Luo, “Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images,” in CVPR . IEEE, 2019, pp. 5337–5345
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , 2020
2020
Earlier work this paper cites.
M. Saito, S. Saito, M. Koyama, and S. Kobayashi, “Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal GAN,” Int. J. Comput. Vis. , vol. 128, no. 10, pp. 2586–2606, 2020
2020
Earlier work this paper cites.
Z. Teed and J. Deng, “RAFT: recurrent all-pairs field transforms for optical flow,” in ECCV , vol. 12347. Springer, 2020, pp. 402–419
2020
Earlier work this paper cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in CVPR . IEEE, 2020, pp. 8107–8116
2020
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR , 2021
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in ICML , vol. 139. PMLR, 2021, pp. 8162–8171
2021
Earlier work this paper cites.
P. Dhariwal and A. Q. Nichol, “Diffusion models beat gans on image synthesis,” in NeurIPS , 2021, pp. 8780–8794
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in CVPR . IEEE, 2021, pp. 12 873–12 883
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Kasten, D. Ofri, O. Wang, and T. Dekel, “Layered neural atlases for consistent video editing,” ACM Trans. Graph. , vol. 40, no. 6, pp. 210:1–210:12, 2021
2021
Earlier work this paper cites.
Z. Su, W. Liu, Z. Yu, D. Hu, Q. Liao, Q. Tian, M. Pietikäinen, and L. Liu, “Pixel difference networks for efficient edge detection,” in ICCV . IEEE, 2021, pp. 5097–5107
2021
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in ICCV . IEEE, 2021, pp. 1708–1718
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , vol. 139. PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
Y. Jafarian and H. S. Park, “Learning high fidelity depths of dressed humans by watching social media dance videos,” in CVPR . IEEE, 2021, pp. 12 753–12 762
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in ICCV . IEEE, 2021, pp. 9630–9640
2021
Earlier work this paper cites.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen, “GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,” in ICML , vol. 162. PMLR, 2022, pp. 16 784–16 804
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, S. K. S. Ghasemipour, R. G. Lopes, B. K. Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi, “Photorealistic text-to-image diffusion models with deep language understanding,” in NeurIPS , 2022
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR . IEEE, 2022, pp. 10 674–10 685
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
G. Kim, T. Kwon, and J. C. Ye, “Diffusionclip: Text-guided diffusion models for robust image manipulation,” in CVPR . IEEE, 2022, pp. 2416–2425
2022
Earlier work this paper cites.
C. Saharia, W. Chan, H. Chang, C. A. Lee, J. Ho, T. Salimans, D. J. Fleet, and M. Norouzi, “Palette: Image-to-image diffusion models,” in SIGGRAPH . ACM, 2022, pp. 15:1–15:10
2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” in NeurIPS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” CoRR , vol. abs/2207.12598, 2022
2022
Earlier work this paper cites.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” in NeurIPS , 2022
2022
Earlier work this paper cites.
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, “Cascaded diffusion models for high fidelity image generation,” J. Mach. Learn. Res. , vol. 23, pp. 47:1–47:33, 2022
2022
Earlier work this paper cites.
S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo, “Vector quantized diffusion model for text-to-image synthesis,” in CVPR . IEEE, 2022, pp. 10 686–10 696
2022
Earlier work this paper cites.
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J. Zhu, and S. Ermon, “Sdedit: Guided image synthesis and editing with stochastic differential equations,” in ICLR . OpenReview.net, 2022
2022
Earlier work this paper cites.
O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in CVPR . IEEE, 2022, pp. 18 187–18 197
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in ICLR , 2022
2022
Earlier work this paper cites.
I. Skorokhodov, S. Tulyakov, and M. Elhoseiny, “Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2,” in CVPR . IEEE, 2022, pp. 3616–3626
2022
Earlier work this paper cites.
V. Voleti, A. Jolicoeur-Martineau, and C. Pal, “MCVD - masked conditional video diffusion for prediction, generation, and interpolation,” in NeurIPS , 2022
2022
Earlier work this paper cites.
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood, “Flexible diffusion modeling of long videos,” in NeurIPS , 2022
2022
Earlier work this paper cites.
A. W. Harley, Z. Fang, and K. Fragkiadaki, “Particle video revisited: Tracking through occlusions using point trajectories,” in ECCV , vol. 13682. Springer, 2022, pp. 59–75
2022
Earlier work this paper cites.
V. K. M. Vadakital, A. Dziembowski, G. Lafruit, F. Thudor, G. Lee, and P. Rondao-Alface, “The MPEG immersive video standard - current status and future outlook,” IEEE Multim. , vol. 29, no. 3, pp. 101–111, 2022
2022
Earlier work this paper cites.
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao, “Gmflow: Learning optical flow via global matching,” in CVPR . IEEE, 2022, pp. 8111–8120
2022
Earlier work this paper cites.
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 3, pp. 1623–1637, 2022
2022
Earlier work this paper cites.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J. Hwang, K. Chang, and J. Gao, “Grounded language-image pre-training,” in CVPR . IEEE, 2022, pp. 10 955–10 965
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. C. H. Hoi, “BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , vol. 162. PMLR, 2022, pp. 12 888–12 900
2022
Cited alongside, same era.
OpenAI, “ChatGPT,” https://openai.com/chatgpt/ , 2022
2022
Cited alongside, same era.
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in ECCV , vol. 13688. Springer, 2022, pp. 640–658
2022
Cited alongside, same era.
D. Watson, W. Chan, J. Ho, and M. Norouzi, “Learning fast samplers for diffusion models by differentiating through sample quality,” in ICLR , 2022
2023
Later among the works it cites.
W. Chai, X. Guo, G. Wang, and Y. Lu, “Stablevideo: Text-driven consistency-aware diffusion video editing,” in ICCV . IEEE, 2023, pp. 22 983–22 993
2023
Later among the works it cites.
Y. Lee, J. G. Jang, Y. Chen, E. Qiu, and J. Huang, “Shape-aware text-driven layered video editing,” in CVPR . IEEE, 2023, pp. 14 317–14 326
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
W. Wang, Q. Lai, H. Fu, J. Shen, H. Ling, and R. Yang, “Salient object detection in the deep learning era: An in-depth survey,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 6, pp. 3239–3259, 2022
2022
Cited alongside, same era.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in CVPR . IEEE, 2022, pp. 1280–1289
2022
Cited alongside, same era.
P. Truong, M. Danelljan, F. Yu, and L. V. Gool, “Probabilistic warp consistency for weakly-supervised semantic correspondences,” in CVPR . IEEE, 2022, pp. 8698–8708
2022
Cited alongside, same era.
J. Fu, S. Li, Y. Jiang, K. Lin, C. Qian, C. C. Loy, W. Wu, and Z. Liu, “Stylegan-human: A data-centric odyssey of human generation,” in ECCV , vol. 13676. Springer, 2022, pp. 1–19
2022
Cited alongside, same era.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,” Trans. Mach. Learn. Res. , vol. 2022, 2022
2022
Cited alongside, same era.
H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo, “Advancing high-resolution video-language representation with large-scale video transcriptions,” in CVPR . IEEE, 2022, pp. 5026–5035
2022
Cited alongside, same era.
B. Lefaudeux, F. Massa, D. Liskovich, W. Xiong, V. Caggiano, S. Naren, M. Xu, J. Hu, M. Tintore, S. Zhang, P. Labatut, D. Haziza, L. Wehrstedt, J. Reizenstein, and G. Sizov, “xFormers: A modular and hackable transformer modelling library,” https://github.com/facebookresearch/xformers , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Yang, Y. Zhou, Z. Liu, and C. C. Loy, “Rerender A video: Zero-shot text-guided video-to-video translation,” in SIGGRAPH . ACM, 2023, pp. 95:1–95:11
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi, “Text2video-zero: Text-to-image diffusion models are zero-shot video generators,” in ICCV . IEEE, 2023, pp. 15 908–15 918
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Liu, Y. Zhang, W. Li, Z. Lin, and J. Jia, “Video-p2p: Video editing with cross-attention control,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Qi, X. Cun, Y. Zhang, C. Lei, X. Wang, Y. Shan, and Q. Chen, “Fatezero: Fusing attentions for zero-shot text-based video editing,” in ICCV . IEEE, 2023, pp. 15 886–15 896
2023
Later among the works it cites.
C. Shin, H. Kim, C. H. Lee, S. Lee, and S. Yoon, “Edit-a-video: Single video editing with object-aware consistency,” in ACML , vol. 222. PMLR, 2023, pp. 1215–1230
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou, “Videocomposer: Compositional video synthesis with motion controllability,” in NeurIPS , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou, “Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,” in ICCV . IEEE, 2023, pp. 7589–7599
2023
Later among the works it cites.
Z. Zhang, B. Li, X. Nie, C. Han, T. Guo, and L. Liu, “Towards consistent video editing with text-to-image diffusion models,” in NeurIPS , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Ge, S. Nah, G. Liu, T. Poon, A. Tao, B. Catanzaro, D. Jacobs, J. Huang, M. Liu, and Y. Balaji, “Preserve your own correlation: A noise prior for video diffusion models,” in ICCV . IEEE, 2023, pp. 22 873–22 884
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Pan, A. Tewari, T. Leimkühler, L. Liu, A. Meka, and C. Theobalt, “Drag your GAN: interactive point-based manipulation on the generative image manifold,” in SIGGRAPH . ACM, 2023, pp. 78:1–78:11
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” in ICCV . IEEE, 2023, pp. 20 087–20 097
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy, “Pick-a-pic: An open dataset of user preferences for text-to-image generation,” in NeurIPS , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency models,” in ICML , vol. 202. PMLR, 2023, pp. 32 211–32 252
2023
Later among the works it cites.
C. Meng, R. Rombach, R. Gao, D. P. Kingma, S. Ermon, J. Ho, and T. Salimans, “On distillation of guided diffusion models,” in CVPR . IEEE, 2023, pp. 14 297–14 306
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV . IEEE, 2023, pp. 4172–4182
2023
Later among the works it cites.
2024
Closest in time.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Comput. Surv. , vol. 56, no. 4, pp. 105:1–105:39, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan, “T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in AAAI . AAAI Press, 2024, pp. 4296–4304
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Ma, Y. He, X. Cun, X. Wang, S. Chen, X. Li, and Q. Chen, “Follow your pose: Pose-guided text-to-video generation using pose-free videos,” in AAAI . AAAI Press, 2024, pp. 4117–4125
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Ling, L. Chen, P. Zhang, H. Chen, Y. Jin, and J. Zheng, “Freedrag: Feature dragging for reliable point-based image editing,” in CVPR . IEEE, 2024, pp. 6860–6870
2024
Closest in time.
OpenAI, “GPT-4o,” https://openai.com/index/hello-gpt-4o/ , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.