Fetching the paper…
Reading the bibliography…
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space.
Coding theorems for a discrete source with a fidelity criterion
Shannon, C. E. et al · 1959
Earlier work this paper cites.
Logistic-normal distributions: Some properties and uses
Atchison, J. and Shen, S. M · 1980
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Improved techniques for training GANs
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
The perception-distortion tradeoff
Blau, Y. and Michaeli, T · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Which training methods for GANs do actually converge?
Mescheder, L., Geiger, A., and Nowozin, S · 2018
Earlier work this paper cites.
FiLM: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Group normalization
Wu, Y. and He, K · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Rethinking lossy compression: The rate-distortion-perception tradeoff
Blau, Y. and Michaeli, T · 2019
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., and Aila, T · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with VQ-VAE-2
Razavi, A., Van den Oord, A., and Vinyals, O · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
PULSE: Self-supervised photo upsampling via latent space exploration of generative models
Menon, S., Damian, A., Hu, S., Ravi, N., and Rudin, C · 2020
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
MAGVIT: Masked generative video transformer
Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., et al · 2023
Later among the works it cites.
Designing a better asymmetric VQGAN for StableDiffusion
Zhu, Z., Feng, X., Chen, D., Bao, J., Wang, L., Chen, Y., Yuan, L., and Hua, G · 2023
Later among the works it cites.
Baldridge, J., Bauer, J., Bhutani, M., Brichtova, N., Bunner, A., Chan, K., Chen, Y., Dieleman, S., Du, Y., Eaton-Rosen, Z., et al · 2024
Closest in time.
Birodkar, V., Barcik, G., Lyon, J., Ioffe, S., Minnen, D., and Dillon, J. V · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Image representations learned with unsupervised pre-training contain human-like biases
Steed, R. and Caliskan, A · 2021
Cited alongside, same era.
MaskGIT: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Cascaded diffusion models for high fidelity image generation
Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T · 2022
Cited alongside, same era.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A · 2024
Closest in time.
Deep compression autoencoder for efficient high-resolution diffusion models
Chen, J., Cai, H., Chen, J., Xie, E., Yang, S., Tang, H., Li, M., Lu, Y., and Han, S · 2024
Closest in time.
Patched denoising diffusion models for high-resolution image synthesis
Ding, Z., Zhang, M., Wu, J., and Tu, Z · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R · 2024
Closest in time.
Flax: A neural network library and ecosystem for JAX, 2024
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2024
Closest in time.
Understanding diffusion objectives as the elbo with simple data augmentation
Kingma, D. and Gao, R · 2024
Closest in time.
VideoPoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2024
Closest in time.
Improving the training of rectified flows
Lee, S., Lin, Z., and Fanti, G · 2024
Closest in time.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Closest in time.
SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S · 2024
Closest in time.
SwiftBrush: One-step text-to-image diffusion model with variational score distillation
Nguyen, T. H. and Tran, A · 2024
Closest in time.
Würstchen: An efficient architecture for large-scale text-to-image diffusion models
Pernias, P., Rampas, D., Richter, M. L., Pal, C. J., and Aubreville, M · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2024
Closest in time.
LiteVAE: Lightweight and efficient variational autoencoders for latent diffusion models
Sadat, S., Buhmann, J., Bradley, D., Hilliges, O., and Weber, R. M · 2024
Closest in time.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R · 2024
Closest in time.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Closest in time.
Lossy image compression with conditional diffusion models
Yang, R. and Mandt, S · 2024
Closest in time.
Image and video tokenization with binary spherical quantization
Zhao, Y., Xiong, Y., and Krähenbühl, P · 2024
Closest in time.
Diffusion autoencoders are scalable image tokenizers
Chen, Y., Girdhar, R., Wang, X., Rambhatla, S. S., and Misra, I · 2025
Closest in time.
Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization
Sargent, K., Hsu, K., Johnson, J., Fei-Fei, L., and Wu, J · 2025
Closest in time.