Fetching the paper…
Reading the bibliography…
We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang, et al · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Coyo-700m: Image-text pair dataset
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Earlier work this paper cites.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang, et al · 2022
Earlier work this paper cites.
Autoregressive image generation using residual quantization
D. Lee, C. Kim, S. Kim, M. Cho, and W.-S. Han · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Earlier work this paper cites.
Movq: Modulating quantized vectors for high-fidelity image generation
C. Zheng, T.-L. Vuong, J. Cai, and D. Phung · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Earlier work this paper cites.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Earlier work this paper cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen · 2023
Earlier work this paper cites.
Openorca: An open dataset of gpt augmented flan reasoning traces
W. Lian, B. Goodson, E. Pentland, A. Cook, C. Vong, and "Teknium" · 2023
Earlier work this paper cites.
On the hidden mystery of ocr in large multimodal models
Y. Liu, Z. Li, B. Yang, C. Li, X. Yin, C.-l. Liu, L. Jin, and X. Bai · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Cited alongside, same era.
Generative pretraining in multimodality
Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang · 2023
Cited alongside, same era.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Teknium · 2023
Cited alongside, same era.
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi · 2024
Later among the works it cites.
Emu edit: Precise image editing via recognition and generation tasks
S. Sheynin, A. Polyak, U. Singer, Y. Kirstain, A. Zohar, O. Ashual, D. Parikh, and Y. Taigman · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
P. Sun, Y. Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
C. Team · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
K. Tian, Y. Jiang, Z. Yuan, B. Peng, and L. Wang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Cited alongside, same era.
Deep compression autoencoder for efficient high-resolution diffusion models
J. Chen, H. Cai, J. Chen, E. Xie, S. Yang, H. Tang, M. Li, Y. Lu, and S. Han · 2024
Cited alongside, same era.
Emova: Empowering language models to see, hear and speak with vivid emotions
K. Chen, Y. Gou, R. Huang, Z. Liu, D. Tan, J. Xu, C. Wang, Y. Zhu, Y. Zeng, K. Yang, et al · 2024
Cited alongside, same era.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Cited alongside, same era.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
P. Tong, E. Brown, P. Wu, S. Woo, A. J. V. IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al · 2024
Later among the works it cites.
Illume: Illuminating your llms to see, draw, and self-enhance
C. Wang, G. Lu, J. Yang, R. Huang, J. Han, L. Hou, W. Zhang, and H. Xu · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Later among the works it cites.
Emu3: Next-token prediction is all you need
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y. Cui, J. Wang, F. Zhang, Y. Wang, Z. Li, Q. Yu, et al · 2024
Later among the works it cites.
Omniedit: Building image editing generalist models through specialist supervision
C. Wei, Z. Xiong, W. Ren, X. Du, G. Zhang, and W. Chen · 2024
Later among the works it cites.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan, et al · 2024
Later among the works it cites.
Vila-u: a unified foundation model integrating visual understanding and generation
Y. Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y. Fang, L. Zhu, E. Xie, H. Yin, L. Yi, et al · 2024
Later among the works it cites.
Omnigen: Unified image generation
S. Xiao, Y. Wang, J. Zhou, H. Yuan, X. Xing, R. Yan, S. Wang, T. Huang, and Z. Liu · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou · 2024
Later among the works it cites.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Z. Xu, F. Jiang, L. Niu, Y. Deng, R. Poovendran, Y. Choi, and B. Y. Lin · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Later among the works it cites.
X-vila: Cross-modality alignment for large language model
H. Ye, D.-A. Huang, Y. Lu, Z. Yu, W. Ping, A. Tao, J. Kautz, S. Han, D. Xu, P. Molchanov, et al · 2024
Later among the works it cites.
Anyedit: Mastering unified high-quality image editing for any idea
Q. Yu, W. Chow, Z. Yue, K. Pan, Y. Wu, X. Wan, J. Li, S. Tang, H. Zhang, and Y. Zhuang · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Later among the works it cites.
Magicbrush: A manually annotated dataset for instruction-guided image editing
K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su · 2024
Later among the works it cites.
Ultraedit: Instruction-based fine-grained image editing at scale
H. Zhao, X. S. Ma, L. Chen, S. Si, R. Wu, K. An, P. Yu, M. Zhang, Q. Li, and B. Chang · 2024
Later among the works it cites.
Addressing representation collapse in vector quantized models with one linear layer
Y. Zhu, B. Li, Y. Xin, and L. Xu · 2024
Later among the works it cites.
Unit: Unifying image and text recognition in one vision encoder
Y. Zhu, Y. Zhou, C. Wang, Y. Cao, J. Han, L. Hou, and H. Xu · 2024
Later among the works it cites.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan · 2025
Closest in time.
Evaluating text-to-visual generation with image-to-text generation
Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan · 2025
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2025
Closest in time.
D. Lu, X. Tan, R. Xu, T. Yao, C. Qu, W. Chu, Y. Xu, and Y. Qi · 2025
Closest in time.