Fetching the paper…
Reading the bibliography…
Recent years have seen significant progress in human image generation, particularly with the advancements in diffusion models.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Learning multiple tasks with multilinear relationship networks
M. Long, Z. Cao, J. Wang, and P. S. Yu · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
A. Kendall, Y. Gal, and R. Cipolla · 2017
Earlier work this paper cites.
Pose guided person image generation
L. Ma, X. Jia, Q. Sun, B. Schiele, T. Tuytelaars, and L. Van Gool · 2017
Earlier work this paper cites.
Demystifying MMD gans
M. Binkowski, D. J. Sutherland, M. Arbel, and A. Gretton · 2018
Earlier work this paper cites.
Deformable gans for pose-based human image generation
A. Siarohin, E. Sangineto, S. Lathuilière, and N. Sebe · 2018
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C. Chang, M. G. Yong, J. Lee, W. Chang, W. Hua, M. Georg, and M. Grundmann · 2019
Earlier work this paper cites.
Progressive pose attention transfer for person image generation
Z. Zhu, T. Huang, B. Shi, M. Yu, B. Wang, and X. Bai · 2019
Earlier work this paper cites.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
C. Zimmermann, D. Ceylan, J. Yang, B. Russel, M. Argus, and T. Brox · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image
G. Moon, S.-I. Yu, H. Wen, T. Shiratori, and K. M. Lee · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Everybody sign now: Translating spoken language to photo realistic sign language video
B. Saunders, N. C. Camgoz, and R. Bowden · 2020
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Cited alongside, same era.
Signing at scale: Learning to co-articulate signs for large-scale photo-realistic sign language production
B. Saunders, N. C. Camgoz, and R. Bowden · 2022
Later among the works it cites.
Cross attention based style distribution for controllable person image synthesis
X. Zhou, M. Yin, X. Chen, L. Sun, C. Gao, and Q. Li · 2022
Later among the works it cites.
Blended latent diffusion
O. Avrahami, O. Fried, and D. Lischinski · 2023
Later among the works it cites.
Multimodal garment designer: Human-centric latent diffusion models for fashion image editing
A. Baldrati, D. Morelli, G. Cartella, M. Cornia, M. Bertini, and R. Cucchiara · 2023
Later among the works it cites.
Person image synthesis via denoising diffusion model
A. K. Bhunia, S. Khan, H. Cholakkal, R. M. Anwer, J. Laaksonen, M. Shah, and F. S. Khan · 2023
Later among the works it cites.
Concept sliders: Lora adaptors for precise control in diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blended diffusion for text-driven editing of natural images
O. Avrahami, D. Lischinski, and O. Fried · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Torchmetrics - measuring reproducibility in pytorch
N. S. Detlefsen, J. Borovec, J. Schock, A. H. Jha, T. Koker, L. D. Liello, D. Stancl, C. Quan, M. Grechkin, and W. Falcon · 2022
Cited alongside, same era.
Hagrid - hand gesture recognition image dataset
A. Kapitanov, A. Makhlyarchuk, and K. Kvanchiani · 2022
Cited alongside, same era.
R. Gandikota, J. Materzyńska, T. Zhou, A. Torralba, and D. Bau · 2023
Later among the works it cites.
Humansd: A native skeleton-guided diffusion model for human image generation
X. Ju, A. Zeng, C. Zhao, J. Wang, L. Zhang, and Q. Xu · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Later among the works it cites.
Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting
W. Lu, Y. Xu, J. Zhang, C. Wang, and D. Tao · 2023
Later among the works it cites.
A dataset of relighted 3D interacting hands
G. Moon, S. Saito, W. Xu, R. Joshi, J. Buffalini, H. Bellan, N. Rosen, J. Richardson, M. Mallorie, P. Bree, T. Simon, B. Peng, S. Garg, K. McPhail, and T. Shiratori · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Later among the works it cites.
Advancing pose-guided image synthesis with progressive conditional diffusion models
F. Shen, H. Ye, J. Zhang, C. Wang, X. Han, and W. Yang · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan · 2024
Closest in time.