Fetching the paper…
Reading the bibliography…
Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
Ranzato, M., Huang, F.-J., Boureau, Y.-L., and LeCun, Y · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Deep Boltzmann machines
Salakhutdinov, R. and Hinton, G · 2009
Earlier work this paper cites.
Stacked convolutional auto-encoders for hierarchical feature extraction
Masci, J., Meier, U., Cires, D., and Schmidhuber, J · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Adversarial feature learning
Donahue, J., Krahenbühl, P., and Darrell, T · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network
Ledig, C., Theis, L., Huszár, F., Caballero, J., Cunningham, A., Acosta, A., Aitken, A., Tejani, A., Totz, J., Wang, Z., et al · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Esrgan: Enhanced super-resolution generative adversarial networks
Wang, X., Yu, K., Wu, S., Gu, J., Liu, Y., Dong, C., Qiao, Y., and Change Loy, C · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Pros and cons of gan evaluation measures
Borji, A · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Self-supervised learning of pretext-invariant representations
Misra, I. and Maaten, L. v. d · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Later among the works it cites.
Align your latents: High-resolution video synthesis with latent diffusion models
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K · 2023
Later among the works it cites.
Emu: Enhancing image generation models using photogenic needles in a haystack
Dai, X., Hou, J., Ma, C.-Y., Tsai, S., Wang, J., Wang, R., Zhang, P., Vandenhende, S., Wang, X., Dubey, A., et al · 2023
Later among the works it cites.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Later among the works it cites.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
Real-esrgan: Training real-world blind super-resolution with pure synthetic data
Wang, X., Xie, L., Dong, C., and Shan, Y · 2021
Cited alongside, same era.
Building normalizing flows with stochastic interpolants
Albergo, M. S. and Vanden-Eijnden, E · 2022
Cited alongside, same era.
BEiT: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2022
Cited alongside, same era.
Pros and cons of gan evaluation measures: New developments
Borji, A · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Consistency models
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Later among the works it cites.
Sample what you cant compress, 2024
Birodkar, V., Barcik, G., Lyon, J., Ioffe, S., Minnen, D., and Dillon, J. V · 2024
Later among the works it cites.
Image neural field diffusion models
Chen, Y., Wang, O., Zhang, R., Shechtman, E., Wang, X., and Gharbi, M · 2024
Later among the works it cites.
Rethinking fid: Towards a better evaluation metric for image generation
Jayasumana, S., Ramalingam, S., Veit, A., Glasner, D., Chakrabarti, A., and Kumar, S · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Understanding diffusion objectives as the elbo with simple data augmentation
Kingma, D. and Gao, R · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Later among the works it cites.
Simplifying, stabilizing and scaling continuous-time consistency models
Lu, C. and Song, Y · 2024
Later among the works it cites.
Würstchen: An efficient architecture for large-scale text-to-image diffusion models
Pernias, P., Rampas, D., Richter, M. L., Pal, C., and Aubreville, M · 2024
Later among the works it cites.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al · 2024
Later among the works it cites.
Multistep distillation of diffusion models via moment matching
Salimans, T., Mensink, T., Heek, J., and Hoogeboom, E · 2024
Later among the works it cites.
Improved techniques for training consistency models
Song, Y. and Dhariwal, P · 2024
Later among the works it cites.
Em distillation for one-step diffusion models
Xie, S., Xiao, Z., Kingma, D. P., Hou, T., Wu, Y. N., Murphy, K. P., Salimans, T., Poole, B., and Gao, R · 2024
Later among the works it cites.
ϵ \epsilon -vae: Denoising as visual decoding, 2024
Zhao, L., Woo, S., Wan, Z., Li, Y., Zhang, H., Gong, B., Adam, H., Jia, X., and Liu, T · 2024
Later among the works it cites.