Fetching the paper…
Reading the bibliography…
We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space.
Interference in visual recognition
Jerome S. Bruner and Mary C. Potter · 1964
Earlier work this paper cites.
Forest before trees: The precedence of global features in visual perception
David Navon · 1977
Earlier work this paper cites.
Face recognition using eigenfaces
Matthew A. Turk and Alex P. Pentland · 1991
Earlier work this paper cites.
Diagnostic colors mediate scene recognition
Aude Oliva and Philippe G. Schyns · 2000
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E. Hinton and Ruslan R. Salakhutdinov · 2006
Earlier work this paper cites.
What do we perceive in a glance of a real-world scene?
Li Fei-Fei, Asha Iyer, Christof Koch, and Pietro Perona · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Learning ordered representations with nested dropout
Oren Rippel, Michael Gelbart, and Ryan Adams · 2014
Earlier work this paper cites.
A tutorial on principal component analysis
Jonathon Shlens · 2014
Earlier work this paper cites.
Improved techniques for training GANs
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Learning deep Transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
GLU variants improve Transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
Prafulla Dhariwal and Alex Nichol · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Taming Transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
MaskGIT: Masked generative image transformer
Diffusion is spectral autoregression, 2024
Sander Dieleman · 2024
Later among the works it cites.
Divot: Diffusion powers video tokenizer for comprehension and generation
Yuying Ge, Yizhuo Li, Yixiao Ge, and Ying Shan · 2024
Later among the works it cites.
SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Nanye Ma, Mark Goldstein, Michael S Albergo, Nicholas M Boffi, Eric Vanden-Eijnden, and Saining Xie · 2024
Later among the works it cites.
DINOv2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Later among the works it cites.
When worse is better: Navigating the compression-generation tradeoff in visual tokenization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, and Ali Farhadi · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Cited alongside, same era.
Diffusion autoencoders: Toward a meaningful and decodable representation
Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, and Supasorn Suwajanakorn · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
Vivek Ramanujan, Kushal Tirumala, Armen Aghajanyan, Luke Zettlemoyer, and Ali Farhadi · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Later among the works it cites.
MaskBit: Embedding-free image generation via bit tokens
Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen · 2024
Later among the works it cites.
ϵ \epsilon -VAE: Denoising as visual decoding
Long Zhao, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang, Boqing Gong, Hartwig Adam, Xuhui Jia, and Ting Liu · 2024
Later among the works it cites.
Scaling the codebook size of VQ-GAN to 100,000 with a utilization rate of 99%
Lei Zhu, Fangyun Wei, Yanye Lu, and Dong Chen · 2024
Later among the works it cites.
FlexTok: Resampling images into 1d token sequences of flexible length
Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oğuzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Zamir, and Afshin Dehghan · 2025
Closest in time.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin · 2025
Closest in time.
Generative modelling in latent space, 2025
Sander Dieleman · 2025
Closest in time.
Adaptive length image tokenization via recurrent allocation
Shivam Duggal, Phillip Isola, Antonio Torralba, and William T. Freeman · 2025
Closest in time.
Learnings from scaling visual tokenizers for reconstruction and generation
Philippe Hansen-Estruch, David Yan, Ching-Yao Chung, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, and Xinlei Chen · 2025
Closest in time.
Imagefolder: Autoregressive image generation with folded tokens
Xiang Li, Kai Qiu, Hao Chen, Jason Kuen, Jiuxiang Gu, Bhiksha Raj, and Zhe Lin · 2025
Closest in time.
One-d-piece: Image tokenizer meets quality-controllable compression
Keita Miwa, Kento Sasaki, Hidehisa Arai, Tsubasa Takahashi, and Yu Yamaguchi · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Jingfeng Yao and Xinggang Wang · 2025
Closest in time.
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie · 2025
Closest in time.