Fetching the paper…
Reading the bibliography…
Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Conditional generative adversarial nets
Mirza, M.; and Osindero, S. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A.; Vinyals, O.; and Kavukcuoglu, K. 2018 · 2018
Earlier work this paper cites.
Video-to-Video Synthesis
Wang, T.-C.; Liu, M.-Y.; Zhu, J.-Y.; Liu, G.; Tao, A.; Kautz, J.; and Catanzaro, B. 2018 · 2018
Earlier work this paper cites.
Everybody Dance Now
Chan, C.; Ginosar, S.; Zhou, T.; and Efros, A. A. 2019 · 2019
Earlier work this paper cites.
Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis
Liu, W.; Piao, Z.; Min, J.; Luo, W.; Ma, L.; and Gao, S. 2019 · 2019
Earlier work this paper cites.
First order motion model for image animation
Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; and Sebe, N. 2019 · 2019
Earlier work this paper cites.
HRNet: Deep High-Resolution Representation Learning for Human Pose Estimation
Sun, K.; Xiao, B.; Liu, D.; and Wang, J. 2019 · 2019
Earlier work this paper cites.
Few-shot Video-to-Video Synthesis
Wang, T.-C.; Liu, M.-Y.; Tao, A.; Liu, G.; Kautz, J.; and Catanzaro, B. 2019 · 2019
Earlier work this paper cites.
OpenMMLab Pose Estimation Toolbox and Benchmark
Contributors, M. 2020 · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
Sdedit: Image synthesis and editing with stochastic differential equations
Meng, C.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; and Ermon, S. 2021 · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Nguyen, H.-L.; Kim, J.; Yeo, H.; Lee, D.; and Yoon, S.-E. 2021 · 2021
Cited alongside, same era.
Benchmark for Compositional Text-to-Image Synthesis
Park, D. H.; Azadi, S.; Liu, X.; Darrell, T.; and Rohrbach, A. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Visual knowledge graph for human action reasoning in videos
Ma, Y.; Wang, Y.; Wu, Y.; Lyu, Z.; Chen, S.; Li, X.; and Qiao, Y. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Make-a-Video: Text-to-Video Generation without Text-Video Data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. 2022 · 2022
Later among the works it cites.
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Wu, J. Z.; Ge, Y.; Wang, X.; Lei, S. W.; Gu, Y.; Hsu, W.; Shan, Y.; Qie, X.; and Shou, M. Z. 2022 · 2022
Later among the works it cites.
Advancing high-resolution video-language representation with large-scale video transcriptions
Xue, H.; Hang, T.; Zeng, Y.; Sun, Y.; Liu, B.; Yang, H.; Fu, J.; and Guo, B. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Motion representations for articulated animation
Siarohin, A.; Woodford, O. J.; Ren, J.; Chai, M.; and Tulyakov, S. 2021 · 2021
Cited alongside, same era.
Avrahami, O.; Fried, O.; and Lischinski, D. 2022 · 2022
Cited alongside, same era.
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
Ding, M.; Zheng, W.; Hong, W.; and Tang, J. 2022 · 2022
Cited alongside, same era.
Latent Video Diffusion Models for High-Fidelity Video Generation with Arbitrary Lengths
He, Y.; Yang, T.; Zhang, Y.; Shan, Y.; and Chen, Q. 2022 · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
Cao, M.; Wang, X.; Qi, Z.; Shan, Y.; Qie, X.; and Zheng, Y. 2023 · 2023
Closest in time.
Structure and Content-Guided Video Synthesis with Diffusion Models
Esser, P.; Chiu, J.; Atighehchian, P.; Granskog, J.; and Germanidis, A. 2023 · 2023
Closest in time.
Backdoor Defense via Adaptively Splitting Poisoned Dataset
Gao, K.; Bai, Y.; Gu, J.; Yang, Y.; and Xia, S.-T. 2023 · 2023
Closest in time.
He, C.; Li, K.; Zhang, Y.; Zhang, Y.; Guo, Z.; Li, X.; Danelljan, M.; and Yu, F. 2023 · 2023
Closest in time.
FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation
Li, R.; Zhao, J.; Zhang, Y.; Su, M.; Ren, Z.; Zhang, H.; Tang, Y.; and Li, X. 2023 · 2023
Closest in time.
Ma, Z.; Jia, G.; and Zhou, B. 2023 · 2023
Closest in time.
Mou, C.; Wang, X.; Xie, L.; Zhang, J.; Qi, Z.; Shan, Y.; and Qie, X. 2023 · 2023
Closest in time.
FateZero: Fusing Attentions for Zero-shot Text-based Video Editing
Qi, C.; Cun, X.; Zhang, Y.; Lei, C.; Wang, X.; Shan, Y.; and Chen, Q. 2023 · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.