Fetching the paper…
Reading the bibliography…
The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses.
No-reference image quality assessment in the spatial domain
Mittal, A.; Moorthy, A. K.; and Bovik, A. C. 2012 · 2012
Earlier work this paper cites.
Making a “Completely Blind” Image Quality Analyzer
Mittal, A.; Soundararajan, R.; and Bovik, A. C. 2013 · 2013
Earlier work this paper cites.
Generative Adversarial Nets , volume 27 of Advances in Neural Information Processing Systems
Goodfellow, I. J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation , book section Chapter 28, 234–241
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Learning to generate long-term future via hierarchical prediction
Villegas, R.; Yang, J.; Zou, Y.; Sohn, S.; Lin, X.; and Lee, H. 2017 Published · 2017
Earlier work this paper cites.
Chan, C.; Ginosar, S.; Zhou, T.; and Efros, A. A. 2018 · 2018
Earlier work this paper cites.
YOLOv3: An Incremental Improvement
Redmon, J.; and Farhadi, A. 2018 · 2018
Earlier work this paper cites.
Wang, T.-C.; Liu, M.-Y.; Zhu, J.-Y.; Liu, G.; Tao, A.; Kautz, J.; and Catanzaro, B. 2018 · 2018
Earlier work this paper cites.
PyTorch: an imperative style, high-performance deep learning library , Article 721
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Köpf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Earlier work this paper cites.
Few-shot video-to-video synthesis , Article 451
Wang, T.-C.; Liu, M.-Y.; Tao, A.; Liu, G.; Kautz, J.; and Catanzaro, B. 2019 · 2019
Earlier work this paper cites.
ControlVideo: Adding Conditional Control for One Shot Text-to-Video Editing
Zhao, M.; Wang, R.; Bao, F.; Li, C.; and Zhu, J. 2023 · 2019
Earlier work this paper cites.
Array programming with NumPy
Harris, C. R.; Millman, K. J.; van der Walt, S. J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N. J.; Kern, R.; Picus, M.; Hoyer, S.; van Kerkwijk, M. H.; Brett, M.; Haldane, A.; Del Rio, J. F.; Wiebe, M.; Peterson, P.; Gerard-Marchant, P.; Sheppard, K.; Reddy, T.; Weckesser, W.; Abbasi, H.; Gohlke, C.; and Oliphant, T. E. 2020 · 2020
Cited alongside, same era.
OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields
Cao, Z.; Hidalgo, G.; Simon, T.; Wei, S. E.; and Sheikh, Y. 2021 · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 Published · 2022
Later among the works it cites.
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2022 · 2022
Later among the works it cites.
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Zhangjie Wu, J.; Ge, Y.; Wang, X.; Lei, W.; Gu, Y.; Hsu, W.; Shan, Y.; Qie, X.; and Shou, M. Z. 2022 · 2022
Later among the works it cites.
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
Couairon, P.; Rambour, C.; Haugeard, J.-E.; and Thome, N. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brooks, T.; Holynski, A.; and Efros, A. A. 2022 · 2022
Cited alongside, same era.
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Prompt-to-Prompt Image Editing with Cross Attention Control
Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022 · 2022
Cited alongside, same era.
Multi-Concept Customization of Text-to-Image Diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J.-Y. 2022 · 2022
Cited alongside, same era.
Liquid Warping GAN With Attention: A Unified Framework for Human Image Synthesis
Liu, W.; Piao, Z.; Tu, Z.; Luo, W.; Ma, L.; and Gao, S. 2022 · 2022
Cited alongside, same era.
DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; and Zhu, J. 2022 · 2022
Cited alongside, same era.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Dollár, P.; and Girshick, R. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Video-P2P: Video Editing with Cross-attention Control
Liu, S.; Zhang, Y.; Li, W.; Lin, Z.; and Jia, J. 2023 · 2023
Closest in time.
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
Ma, Y.; He, Y.; Cun, X.; Wang, X.; Shan, Y.; Li, X.; and Chen, Q. 2023 · 2023
Closest in time.
Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; Shan, Y.; and Qie, X. 2023 · 2023
Closest in time.
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
Qin, B.; Li, J.; Tang, S.; Chua, T.-S.; and Zhuang, Y. 2023 · 2023
Closest in time.
Interactive Data Synthesis for Systematic Vision Adaptation via LLMs-AIGCs Collaboration
Yu, Q.; Li, J.; Ye, W.; Tang, S.; and Zhuang, Y. 2023 · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.