Fetching the paper…
Reading the bibliography…
This work presents SimpleAR, a vanilla autoregressive visual generation framework without complex architecure modifications.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning, 2020
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Accelerating feedforward computation via parallel nonlinear equation solving
Y. Song, C. Meng, R. Liao, and S. Ermon · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Earlier work this paper cites.
synthetic-dataset-1m-dalle3-high-quality-captions
ProGamerGov · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Earlier work this paper cites.
Training diffusion models with reinforcement learning
K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper · 2023
Cited alongside, same era.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. Kwok, P. Luo, H. Lu, et al · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Aligning text-to-image models using human feedback
K. Lee, H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu · 2023
R. Tian, Q. Dai, J. Bao, K. Qiu, Y. Yang, C. Luo, Z. Wu, and Y.-G. Jiang · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2024
Later among the works it cites.
Omnitokenizer: A joint image-video tokenizer for visual generation
J. Wang, Y. Jiang, Z. Yuan, B. Peng, Z. Wu, and Y.-G. Jiang · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean · 2023
Cited alongside, same era.
Journeydb: A benchmark for generative image understanding
K. Sun, J. Pan, Y. Ge, H. Li, H. Duan, X. Wu, R. Zhang, A. Zhou, Z. Qin, Y. Wang, et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
X. Wu, Y. Hao, K. Sun, Y. Chen, F. Zhu, R. Zhao, and H. Li · 2023
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2024
Cited alongside, same era.
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y. Cui, J. Wang, F. Zhang, Y. Wang, Z. Li, Q. Yu, et al · 2024
Later among the works it cites.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan, et al · 2024
Later among the works it cites.
Liquid: Language models are scalable multi-modal generators
J. Wu, Y. Jiang, C. Ma, Y. Liu, H. Zhao, Z. Yuan, S. Bai, and X. Bai · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2024
Later among the works it cites.
Goku: Flow based video generative foundation models
S. Chen, C. Ge, Y. Zhang, Y. Zhang, F. Zhu, H. Yang, H. Hao, H. Wu, Z. Lai, Y. Hu, et al · 2025
Closest in time.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan · 2025
Closest in time.
Cosmos world foundation model platform for physical ai
N. et. al · 2025
Closest in time.
Seedream 2.0: A native chinese-english bilingual image generation foundation model
L. Gong, X. Hou, F. Li, L. Li, X. Lian, F. Liu, L. Liu, W. Liu, W. Lu, Y. Shi, et al · 2025
Closest in time.
Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis
J. Han, J. Liu, Y. Jiang, B. Yan, Y. Zhang, Z. Yuan, B. Peng, and X. Liu · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Y. Hu, and S. Lin · 2025
Closest in time.
Unitoken: Harmonizing multimodal understanding and generation through unified visual encoding
Y. Jiao, H. Qiu, Z. Jie, S. Chen, J. Chen, L. Ma, and Y.-G. Jiang · 2025
Closest in time.
Unitok: A unified tokenizer for visual generation and understanding
C. Ma, Y. Jiang, J. Wu, J. Yang, X. Yu, Z. Yuan, B. Peng, and X. Qi · 2025
Closest in time.
Movie gen: A cast of media foundation models
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al · 2025
Closest in time.
Accelerating auto-regressive text-to-image generation with training-free speculative jacobi decoding
Y. Teng, H. Shi, X. Liu, X. Ning, G. Dai, Y. Wang, Z. Li, and X. Liu · 2025
Closest in time.
Larp: Tokenizing videos with a learned autoregressive generative prior
H. Wang, S. Suri, Y. Ren, H. Chen, and A. Shrivastava · 2025
Closest in time.
Ddt: Decoupled diffusion transformer
S. Wang, Z. Tian, W. Huang, and L. Wang · 2025
Closest in time.
Vila-u: a unified foundation model integrating visual understanding and generation
Y. Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y. Fang, L. Zhu, E. Xie, H. Yin, L. Yi, et al · 2025
Closest in time.
Show-o: One single transformer to unify multimodal understanding and generation
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou · 2025
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy · 2025
Closest in time.