Fetching the paper…
Reading the bibliography…
Recent image generation schemes typically capture image distribution in a pre-constructed latent space relying on a frozen image tokenizer.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection. arxiv 2020
X Zhu, W Su, L Lu, B Li, X Wang, and J Dai · 2010
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
Alex M Lamb, Anirudh Goyal ALIAS PARTH GOYAL, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
Aäron Van Den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
Prafulla Dhariwal and Alex Nichol · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models, 2021
Alex Nichol and Prafulla Dhariwal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Score-based generative modeling in latent space, 2021
Arash Vahdat, Karsten Kreis, and Jan Kautz · 2021
Cited alongside, same era.
Max-deeplab: End-to-end panoptic segmentation with mask transformers, 2021
Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer, 2022
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2022
Mars: Mixture of auto-regressive models for fine-grained text-to-image synthesis
Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang, Lei Zhang, Zhelun Yu, Haoyuan Li, et al · 2024
Later among the works it cites.
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan · 2024
Later among the works it cites.
4m: Massively multimodal masked modeling
David Mizrahi, Roman Bachmann, Oguzhan Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, and Amir Zamir · 2024
Later among the works it cites.
Efficient autoregressive audio modeling via next-scale prediction
Kai Qiu, Xiang Li, Hao Chen, Jie Sun, Jinglu Wang, Zhe Lin, Marios Savvides, and Bhiksha Raj · 2024
Later among the works it cites.
Flowar: Scale-wise autoregressive image generation meets flow matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Divae: Photorealistic images synthesis with denoising diffusion decoder, 2022
Jie Shi, Chenfei Wu, Jian Liang, Xiang Liu, and Nan Duan · 2022
Cited alongside, same era.
Denoising diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2022
Cited alongside, same era.
Movq: Modulating quantized vectors for high-fidelity image generation, 2022
Chuanxia Zheng, Long Tung Vuong, Jianfei Cai, and Dinh Phung · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Vision transformers need registers, 2023
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2023
Cited alongside, same era.
Peco: Perceptual codebook for bert pre-training of vision transformers
Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, Nenghai Yu, and Baining Guo · 2023
Cited alongside, same era.
Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen · 2024
Later among the works it cites.
Taming scalable visual tokenizer for autoregressive image generation
Fengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang, Ying Shan, and Limin Wang · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Later among the works it cites.
Metamorph: Multimodal understanding and generation via instruction tuning
Shengbang Tong, David Fan, Jiachen Zhu, Yunyang Xiong, Xinlei Chen, Koustuv Sinha, Michael Rabbat, Yann LeCun, Saining Xie, and Zhuang Liu · 2024
Later among the works it cites.
Givt: Generative infinite-vocabulary transformers
Michael Tschannen, Cian Eastwood, and Fabian Mentzer · 2024
Later among the works it cites.
Parallelized autoregressive visual generation
Yuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han, Haoyuan Guo, Zhenheng Yang, Difan Zou, Jiashi Feng, and Xihui Liu · 2024
Later among the works it cites.
Maskbit: Embedding-free image generation via bit tokens
Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen · 2024
Later among the works it cites.
Liquid: Language models are scalable multi-modal generators
Junfeng Wu, Yi Jiang, Chuofan Ma, Yuliang Liu, Hengshuang Zhao, Zehuan Yuan, Song Bai, and Xiang Bai · 2024
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Language-guided image tokenization for generation
Kaiwen Zha, Lijun Yu, Alireza Fathi, David A Ross, Cordelia Schmid, Dina Katabi, and Xiuye Gu · 2024
Later among the works it cites.
Image and video tokenization with binary spherical quantization
Yue Zhao, Yuanjun Xiong, and Philipp Krähenbühl · 2024
Later among the works it cites.
Flextok: Resampling images into 1d token sequences of flexible length
Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oğuzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Zamir, and Afshin Dehghan · 2025
Closest in time.
Masked autoencoders are effective tokenizers for diffusion models
Hao Chen, Yujin Han, Fangyi Chen, Xiang Li, Yidong Wang, Jindong Wang, Ze Wang, Zicheng Liu, Difan Zou, and Bhiksha Raj · 2025
Closest in time.
Democratizing text-to-image masked generative models with compact text-aware one-dimensional tokens
Dongwon Kim, Ju He, Qihang Yu, Chenglin Yang, Xiaohui Shen, Suha Kwak, and Liang-Chieh Chen · 2025
Closest in time.
One-d-piece: Image tokenizer meets quality-controllable compression
Keita Miwa, Kento Sasaki, Hidehisa Arai, Tsubasa Takahashi, and Yu Yamaguchi · 2025
Closest in time.
Beyond next-token: Next-x prediction for autoregressive visual generation
Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Jingfeng Yao and Xinggang Wang · 2025
Closest in time.