Fetching the paper…
Reading the bibliography…
Classifier-Free Guidance (CFG) has been a default technique in various visual generative models, yet it requires inference from both conditional and unconditional models during sampling.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Openclip, July 2021
Ilharco, G., Wortsman, M., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Earlier work this paper cites.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Maximum likelihood training of score-based diffusion models
Song, Y., Durkan, C., Murray, I., and Ermon, S · 2021
Earlier work this paper cites.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
Diffusion posterior sampling for general noisy inverse problems
Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Refining generative process with discriminator guidance in score-based diffusion models
Kim, D., Kim, Y., Kwon, S. J., Kang, W., and Moon, I.-C · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Cited alongside, same era.
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Cited alongside, same era.
On aliased resizing and surprising subtleties in gan evaluation
Parmar, G., Zhang, R., and Zhu, J.-Y · 2022
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Cfg++: Manifold-constrained classifier free guidance for diffusion models
Chung, H., Kim, J., Park, G. Y., Nam, H., and Ye, J. C · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R · 2024
Later among the works it cites.
Fluid: Scaling autoregressive text-to-image generative models with continuous tokens
Fan, L., Li, T., Qin, S., Li, Y., Sun, C., Rubinstein, M., Sun, D., He, K., and Tian, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Egsde: Unpaired image-to-image translation via energy-guided stochastic differential equations
Zhao, M., Bao, F., Li, C., and Zhu, J · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Training diffusion models with reinforcement learning
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2023
Cited alongside, same era.
Later among the works it cites.
Geng, Z., Pokle, A., Luo, W., Lin, J., and Kolter, J. Z · 2024
Later among the works it cites.
Guiding a diffusion model with a bad version of itself
Karras, T., Aittala, M., Kynkäänniemi, T., Lehtinen, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Dynamic negative guidance of diffusion models: Towards immediate content removal
Koulischer, F., Deleu, J., Raya, G., Demeester, T., and Ambrogioni, L · 2024
Later among the works it cites.
Applying guidance in a limited interval improves sample and distribution quality in diffusion models
Kynkäänniemi, T., Aittala, M., Karras, T., Laine, S., Aila, T., and Lehtinen, J · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Later among the works it cites.
Sdxl-lightning: Progressive adversarial diffusion distillation
Lin, S., Wang, A., and Yang, X · 2024
Later among the works it cites.
Simplifying, stabilizing and scaling continuous-time consistency models
Lu, C. and Song, Y · 2024
Later among the works it cites.
Star: Scale-wise text-to-image generation via auto-regressive representations
Ma, X., Zhou, M., Liang, T., Bai, Y., Zhao, T., Chen, H., and Jin, Y · 2024
Later among the works it cites.
Gradient-free classifier guidance for diffusion model sampling
Shenoy, R., Pan, Z., Balakrishnan, K., Cheng, Q., Jeon, Y., Yang, H., and Kim, J · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Later among the works it cites.
Hart: Efficient visual generation with hybrid autoregressive transformer
Tang, H., Wu, Y., Yang, S., Xie, E., Chen, J., Chen, J., Zhang, Z., Cai, H., Lu, Y., and Han, S · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Team, C · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
One-step diffusion with distribution matching distillation
Yin, T., Gharbi, M., Zhang, R., Shechtman, E., Durand, F., Freeman, W. T., and Park, T · 2024
Later among the works it cites.
An image is worth 32 tokens for reconstruction and generation
Yu, Q., Weber, M., Deng, X., Shen, X., Cremers, D., and Chen, L.-C · 2024
Later among the works it cites.
Var-clip: Text-to-image generator with visual auto-regressive modeling
Zhang, Q., Dai, X., Yang, N., An, X., Feng, Z., and Ren, X · 2024
Later among the works it cites.