Fetching the paper…
Reading the bibliography…
Multimodal generative models that can understand and generate across multiple modalities are dominated by autoregressive (AR) approaches, which process tokens sequentially from left to right, or top to bottom.
Clevr-ref+: Diagnosing visual reasoning with referring expressions, 2019
R. Liu, C. Liu, Y. Bai, and A. Yuille · 1901
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
J. Ho, A. Jain, and P. Abbeel · 2006
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server, 2015
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization, Jan. 2019
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations, 2020
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. v. d. Berg · 2021
Earlier work this paper cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang, and J. Tang · 2021
Earlier work this paper cites.
Argmax flows and multinomial diffusion: Learning categorical distributions
E. Hoogeboom, D. Nielsen, P. Jaini, P. Forré, and M. Welling · 2021
Earlier work this paper cites.
Perceiver io: A general architecture for structured inputs & outputs
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Earlier work this paper cites.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Earlier work this paper cites.
Diffusion-lm improves controllable text generation, 2022
X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto · 2022
Earlier work this paper cites.
Clevr-math: A dataset for compositional language, visual, and mathematical reasoning, 2022
A. D. Lindström and S. S. Abraham · 2022
Cited alongside, same era.
Unified-io: A unified model for vision, language, and multi-modal tasks
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality, 2022
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Cited alongside, same era.
Muse: Text-to-image generation via masked generative transformers, 2023
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, Y. Li, and D. Krishnan · 2023
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Later among the works it cites.
Discrete Flow Matching, Nov. 2024
I. Gat, T. Remez, N. Shaul, F. Kreuk, R. T. Q. Chen, G. Synnaeve, Y. Adi, and Y. Lipman · 2024
Later among the works it cites.
Likelihood-based diffusion language models
I. Gulrajani and T. B. Hashimoto · 2024
Later among the works it cites.
Intriguing properties of generative classifiers, 2024
P. Jaini, K. Clark, and R. Geirhos · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
T. Dao · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. H. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence · 2023
Cited alongside, same era.
Geneval: An object-focused framework for evaluating text-to-image alignment, 2023
D. Ghosh, H. Hajishirzi, and L. Schmidt · 2023
Cited alongside, same era.
Efficient diffusion training via min-snr weighting strategy
T. Hang, S. Gu, C. Li, J. Bao, D. Chen, H. Hu, X. Geng, and B. Guo · 2023
Cited alongside, same era.
Unified discrete diffusion for simultaneous vision-language generation
M. Hu, C. Zheng, Z. Yang, T.-J. Cham, H. Zheng, C. Wang, D. Tao, and P. N. Suganthan · 2023
Cited alongside, same era.
Your diffusion model is secretly a zero-shot classifier, 2023
A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, and D. Pathak · 2023
Cited alongside, same era.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
D. Liu, S. Zhao, L. Zhuo, W. Lin, Y. Qiao, H. Li, and P. Gao · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution, 2024
A. Lou, C. Meng, and S. Ermon · 2024
Later among the works it cites.
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Z. Luo, F. Shi, Y. Ge, Y. Yang, L. Wang, and Y. Shan · 2024
Later among the works it cites.
Openelm: An efficient language model family with open-source training and inference framework
S. Mehta, M. H. Sekhavat, Q. Cao, M. Horton, Y. Jin, C. Sun, I. Mirzadeh, M. Najibi, D. Belenko, P. Zatloukal, et al · 2024
Later among the works it cites.
Scaling up masked diffusion models on text, 2024
S. Nie, F. Zhu, C. Du, T. Pang, Q. Liu, G. Zeng, M. Lin, and C. Li · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V. Kuleshov · 2024
Later among the works it cites.
Stretching each dollar: Diffusion training from scratch on a micro-budget
V. Sehwag, X. Kong, J. Li, M. Spranger, and L. Lyu · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
J. Shi, K. Han, Z. Wang, A. Doucet, and M. K. Titsias · 2024
Later among the works it cites.
From pixels to prose: A large dataset of dense image captions, 2024
V. Singla, K. Yue, S. Paul, R. Shirkavand, M. Jayawardhana, A. Ganjdanesh, H. Huang, A. Bhatele, G. Somepalli, and T. Goldstein · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
P. Sun, Y. Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan · 2024
Later among the works it cites.
Any-to-any generation via composable diffusion
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, et al · 2024
Later among the works it cites.
Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling, 2024
K. Zheng, Y. Chen, H. Mao, M.-Y. Liu, J. Zhu, and Q. Zhang · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy · 2024
Later among the works it cites.
Lumina-next: Making lumina-t2x stronger and faster with next-dit
L. Zhuo, R. Du, H. Xiao, Y. Li, D. Liu, R. Huang, W. Liu, L. Zhao, F.-Y. Wang, Z. Ma, et al · 2024
Later among the works it cites.
Masked Audio Generation using a Single Non-Autoregressive Transformer, Mar. 2024
A. Ziv, I. Gat, G. L. Lan, T. Remez, F. Kreuk, A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2024
Later among the works it cites.