Fetching the paper…
Reading the bibliography…
Image tokenizers map images to sequences of discrete tokens, and are a crucial component of autoregressive transformer-based image generation.
Biorthogonal bases of compactly supported wavelets
A. Cohen, Ingrid Daubechies, and J.-C. Feauveau · 1992
Earlier work this paper cites.
Ten lectures on wavelets
Ingrid Daubechies · 1992
Earlier work this paper cites.
Jpeg2000: the new still picture compression standard
C. A. Christopoulos, T. Ebrahimi, and A. N. Skodras · 2000
Earlier work this paper cites.
A Wavelet Tour of Signal Processing, Third Edition: The Sparse Way
Stephane Mallat · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, Xi Chen, and Xi Chen · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
PixelSNAIL: An improved autoregressive generative model
XI Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel · 2018
Earlier work this paper cites.
Progressive growing of GANs for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Multi-Stage Variational Auto-Encoders for Coarse-to-Fine Image Generation , page 630–638
Lei Cai, Hongyang Gao, and Shuiwang Ji · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
Singan: Learning a generative model from a single natural image
Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli · 2019
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alex Nichol · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Björn Ommer · 2021
Cited alongside, same era.
Swagan: a style-based wavelet-driven generative model
Rinon Gal, Dana Cohen Hochberg, Amit Bermano, and Daniel Cohen-Or · 2021
Cited alongside, same era.
Generating images with sparse representations
Charlie Nash, Jacob Menick, Sander Dieleman, and Peter W. Battaglia · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Wavelet diffusion models are fast and scalable image generators
Hao Phung, Quan Dao, and Anh Tran · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team · 2023
Later among the works it cites.
Diffusion is spectral autoregression, 2024
Sander Dieleman · 2024
Closest in time.
Quantised global autoencoder: A holistic approach to representing visual data, 2024
Tim Elsner, Paula Usinger, Victor Czech, Gregor Kobsik, Yanjiang He, Isaak Lim, and Leif Kobbelt · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
Matryoshka diffusion models
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Joshua M. Susskind, and Navdeep Jaitly · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman · 2022
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model, 2022
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2022
Cited alongside, same era.
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Cited alongside, same era.
Pyramidal denoising diffusion probabilistic models, 2022
Dohoon Ryu and Jong Chul Ye · 2022
Cited alongside, same era.
Closest in time.
Imagen 3, 2024
Imagen-Team-Google · 2024
Closest in time.
Star: Scale-wise text-to-image generation via auto-regressive representations, 2024
Xiaoxiao Ma, Mohan Zhou, Tao Liang, Yalong Bai, Tiejun Zhao, Huaian Chen, and Yi Jin · 2024
Closest in time.
Wavelets are all you need for autoregressive image generation, 2024
Wael Mattar, Idan Levy, Nir Sharon, and Shai Dekel · 2024
Closest in time.
When worse is better: Navigating the compression-generation tradeoff in visual tokenization, 2024
Vivek Ramanujan, Kushal Tirumala, Armen Aghajanyan, Luke Zettlemoyer, and Ali Farhadi · 2024
Closest in time.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Closest in time.
The llama 3 herd of models, 2024
Llama team · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Closest in time.
Wavelet-based image tokenizer for vision transformers, 2024
Zhenhai Zhu and Radu Soricut · 2024
Closest in time.
Learnings from scaling visual tokenizers for reconstruction and generation, 2025
Philippe Hansen-Estruch, David Yan, Ching-Yao Chung, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, and Xinlei Chen · 2025
Closest in time.
”principal components” enable a new language of images, 2025
Xin Wen, Bingchen Zhao, Ismail Elezi, Jiankang Deng, and Xiaojuan Qi · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Jingfeng Yao, Bin Yang, and Xinggang Wang · 2025
Closest in time.
Frequency autoregressive image generation with continuous tokens, 2025
Hu Yu, Hao Luo, Hangjie Yuan, Yu Rong, and Feng Zhao · 2025
Closest in time.