Fetching the paper…
Reading the bibliography…
Image tokenization has enabled major advances in autoregressive image generation by providing compressed, discrete representations that are more efficient to process than raw pixels.
The jpeg still picture compression standard
Wallace, G. K · 1992
Earlier work this paper cites.
JPEG 2000: Image compression fundamentals, standards and practice
Taubman, D. S. and Marcellin, M. W · 2001
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Learning ordered representations with nested dropout
Rippel, O., Gelbart, M. A., and Adams, R. P · 2014
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2014
Earlier work this paper cites.
Generating images with perceptual similarity metrics based on deep networks
Dosovitskiy, A. and Brox, T · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Baevski, A., Schneider, S., and Auli, M · 2019
Earlier work this paper cites.
Bfloat16 processing for neural networks
Burgess, N., Milanovic, J., Stephens, N., Monachopoulos, K., and Mansell, D. H · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A., Van den Oord, A., and Vinyals, O · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Generative pretraining from pixels
Chen, M., Radford, A., Wu, J., Jun, H., Dhariwal, P., Luan, D., and Sutskever, I · 2020
Earlier work this paper cites.
High-fidelity generative image compression
Mentzer, F., Toderici, G., Tschannen, M., and Agustsson, E · 2020
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N. M · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Earlier work this paper cites.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
Image compression with product quantized masked image modeling
El-Nouby, A., Muckley, M., Ullrich, K., Laptev, I., Verbeek, J., and Jégou, H · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J · 2022
Earlier work this paper cites.
Uvim: A unified modeling approach for vision with learned guiding codes
Kolesnikov, A., Pinto, A. S., Beyer, L., Zhai, X., Harmsen, J., and Houlsby, N · 2022
Cited alongside, same era.
Matryoshka representation learning
Kusupati, A., Bhatt, G., Rege, A., Wallingford, M., Sinha, A., Ramanujan, V., Howard-Snyder, W., Chen, K., Kakade, S. M., Jain, P., and Farhadi, A · 2022
Cited alongside, same era.
Mage: Masked generative encoder to unify representation learning and image synthesis
Li, T., Chang, H., Mishra, S. K., Zhang, H., Katabi, D., and Krishnan, D · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
Unified-io: A unified model for vision, language, and multi-modal tasks
Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A · 2022
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Birodkar, V., Gupta, A., Gu, X., et al · 2023
Later among the works it cites.
Regularized vector quantization for tokenized image synthesis
Zhang, J., Zhan, F., Theobalt, C., and Lu, S · 2023
Later among the works it cites.
4M-21: An any-to-any vision model for tens of tasks and modalities
Bachmann, R., Kar, O. F., Mizrahi, D., Garjani, A., Gao, M., Griffiths, D., Hu, J., Dehghan, A., and Zamir, A · 2024
Later among the works it cites.
Cai, M., Yang, J., Gao, J., and Lee, Y. J · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon, T · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Divae: Photorealistic images synthesis with denoising diffusion decoder
Shi, J., Wu, C., Liang, J., Liu, X., and Duan, N · 2022
Cited alongside, same era.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Cited alongside, same era.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, J. E., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J. W., Chen, W., and Gao, J · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Cited alongside, same era.
Muse: Text-to-image generation via masked generative transformers
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M., Murphy, K. P., Freeman, W. T., Rubinstein, M., Li, Y., and Krishnan, D · 2023
Cited alongside, same era.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Later among the works it cites.
Adaptive length image tokenization via recurrent allocation
Duggal, S., Isola, P., Torralba, A., and Freeman, W. T · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R · 2024
Later among the works it cites.
Multimodal autoregressive pre-training of large vision encoders
Fini, E., Shukor, M., Li, X., Dufter, P., Klein, M., Haldimann, D., Aitharaju, S., da Costa, V. G. T., Béthune, L., Gan, Z., et al · 2024
Later among the works it cites.
Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis
Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., and Liu, X · 2024
Later among the works it cites.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S · 2024
Later among the works it cites.
Midjourney version 6.1 update, 2024
Midjourney · 2024
Later among the works it cites.
Resolving discrepancies in compute-optimal scaling of language models
Porian, T., Wortsman, M., Jitsev, J., Schmidt, L., and Carmon, Y · 2024
Later among the works it cites.
FlexAttention: The Flexibility of PyTorch with the Performance of FlashAttention, August 2024
PyTorch Team: Horace He, Driss Guessous, Yanbo Liang, Joy Dong · 2024
Later among the works it cites.
Eliminating oversaturation and artifacts of high guidance scales in diffusion models
Sadat, S., Hilliges, O., and Weber, R. M · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
How to set adamw’s weight decay as you scale model and dataset size
Wang, X. and Aitchison, L · 2024
Later among the works it cites.
Disco-diff: Enhancing continuous diffusion models with discrete latents
Xu, Y., Corso, G., Jaakkola, T., Vahdat, A., and Kreis, K · 2024
Later among the works it cites.
Elastictok: Adaptive tokenization for image and video
Yan, W., Zaharia, M., Mnih, V., Abbeel, P., Faust, A., and Liu, H · 2024
Later among the works it cites.
Webp image format
Zern, J., Massimino, P., and Alakuijala, J · 2024
Later among the works it cites.
Language-guided image tokenization for generation
Zha, K., Yu, L., Fathi, A., Ross, D. A., Schmid, C., Katabi, D., and Gu, X · 2024
Later among the works it cites.
ϵ \epsilon -vae: Denoising as visual decoding
Zhao, L., Woo, S., Wan, Z., Li, Y., Zhang, H., Gong, B., Adam, H., Jia, X., and Liu, T · 2024
Later among the works it cites.
One-d-piece: Image tokenizer meets quality-controllable compression
Miwa, K., Sasaki, K., Arai, H., Takahashi, T., and Yamaguchi, Y · 2025
Closest in time.
Cat: Content-adaptive image tokenization
Shen, J., Tirumala, K., Yasunaga, M., Misra, I., Zettlemoyer, L., Yu, L., and Zhou, C · 2025
Closest in time.