Fetching the paper…
Reading the bibliography…
Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly interleaved text-image data.
Auto-encoding variational bayes
Kingma, D. P.; Welling, M.; et al. 2013 · 2013
Earlier work this paper cites.
Deep Learning Face Attributes in the Wild
Liu, Z.; Luo, P.; Wang, X.; and Tang, X. 2015 · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
DiffusionDB: A Large-Scale Prompt Gallery Dataset for Text-to-Image Generative Models
Wang, Z. J.; Montoya, E.; Munechika, D.; Yang, H.; Hoover, B.; and Chau, D. H. 2022 · 2022
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T.; Holynski, A.; and Efros, A. A. 2023 · 2023
Earlier work this paper cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Ghosh, D.; Hajishirzi, H.; and Schmidt, L. 2023 · 2023
Earlier work this paper cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
Li, B.; Wang, R.; Wang, G.; Ge, Y.; Ge, Y.; and Shan, Y. 2023 · 2023
Earlier work this paper cites.
JourneyDB: A Benchmark for Generative Image Understanding
Pan, J.; Sun, K.; Ge, Y.; Li, H.; Duan, H.; Wu, X.; Zhang, R.; Zhou, A.; Qin, Z.; Wang, Y.; Dai, J.; Qiao, Y.; and Li, H. 2023 · 2023
Earlier work this paper cites.
Anytext: Multilingual visual text generation and editing
Tuo, Y.; Xiang, W.; He, J.-Y.; Geng, Y.; and Xie, X. 2023 · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Earlier work this paper cites.
Magicbrush: A manually annotated dataset for instruction-guided image editing
Zhang, K.; Mo, L.; Chen, W.; Sun, H.; and Su, Y. 2023 · 2023
Earlier work this paper cites.
FLUX-Controlnet-Inpainting
Creative, A. 2024 · 2024
Earlier work this paper cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; and Ji, R. 2024 · 2024
Cited alongside, same era.
Seed-x: Multimodal models with unified multi-granularity comprehension and generation
Ge, Y.; Zhao, S.; Zhu, J.; Ge, Y.; Yi, K.; Song, L.; Li, C.; Ding, X.; and Shan, Y. 2024 · 2024
Cited alongside, same era.
StyleBooth: Image Style Editing with Multimodal Instruction
Han, Z.; Mao, C.; Jiang, Z.; Pan, Y.; and Zhang, J. 2024 · 2024
Cited alongside, same era.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; Wei, C.; Yu, B.; Yuan, R.; Sun, R.; Yin, M.; Zheng, B.; Yang, Z.; Liu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2024 · 2024
Later among the works it cites.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025 · 2025
Closest in time.
flux-laion-aes
gogoduan. 2025 · 2025
Closest in time.
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
Gu, T.; Yang, K.; Feng, Z.; Wang, X.; Zhang, Y.; Long, D.; Chen, Y.; Cai, W.; and Deng, J. 2025 · 2025
Closest in time.
InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Jiang, L.; Yan, Q.; Jia, Y.; Liu, Z.; Kang, H.; and Lu, X. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hui, M.; Yang, S.; Zhao, B.; Shi, Y.; Wang, H.; Wang, P.; Zhou, Y.; and Xie, C. 2024 · 2024
Cited alongside, same era.
text-to-image-2M
jackyhate. 2024 · 2024
Cited alongside, same era.
laion-high-resolution
LAION. 2024 · 2024
Cited alongside, same era.
Autoregressive image generation without vector quantization
Li, T.; Tian, Y.; Li, H.; Deng, M.; and He, K. 2024 · 2024
Cited alongside, same era.
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2024 · 2024
Cited alongside, same era.
LlamaFusion: Adapting Pretrained Language Models for Multimodal Generation
Shi, W.; Han, X.; Zhou, C.; Liang, W.; Lin, X. V.; Zettlemoyer, L.; and Yu, L. 2024 · 2024
Cited alongside, same era.
Chameleon: Mixed-modal early-fusion foundation models
Team, C. 2024 · 2024
Cited alongside, same era.
Emu3: Next-Token Prediction is All You Need
Wang, X.; Zhang, X.; Luo, Z.; Sun, Q.; Cui, Y.; Wang, J.; Zhang, F.; Wang, Y.; Li, Z.; Yu, Q.; et al. 2024 · 2024
Cited alongside, same era.
Step1x-edit: A practical framework for general image editing
Liu, S.; Han, Y.; Xing, P.; Yin, F.; Wang, R.; Cheng, W.; Liao, J.; Wang, Y.; Fu, H.; Han, C.; et al. 2025 · 2025
Closest in time.
Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Ma, Y.; Liu, X.; Chen, X.; Liu, W.; Wu, C.; Wu, Z.; Pan, Z.; Xie, Z.; Zhang, H.; Yu, X.; et al. 2025 · 2025
Closest in time.
Diffsynth-Studio
ModelScope. 2025 · 2025
Closest in time.
Introducing 4o Image Generation
OpenAI. 2025 · 2025
Closest in time.
Transfer between modalities with metaqueries
Pan, X.; Shukla, S. N.; Singh, A.; Zhao, Z.; Mishra, S. K.; Wang, J.; Xu, Z.; Chen, J.; Li, K.; Juefei-Xu, F.; et al. 2025 · 2025
Closest in time.
Tokenflow: Unified image tokenizer for multimodal understanding and generation
Qu, L.; Zhang, H.; Liu, Y.; Wang, X.; Jiang, Y.; Gao, Y.; Ye, H.; Du, D. K.; Yuan, Z.; and Wu, X. 2025 · 2025
Closest in time.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
Wu, C.; Chen, X.; Wu, Z.; Ma, Y.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C.; et al. 2025 · 2025
Closest in time.
Omnigen: Unified image generation
Xiao, S.; Wang, Y.; Zhou, J.; Yuan, H.; Xing, X.; Yan, R.; Li, C.; Wang, S.; Huang, T.; and Liu, Z. 2025 · 2025
Closest in time.
Anyedit: Mastering unified high-quality image editing for any idea
Yu, Q.; Chow, W.; Yue, Z.; Pan, K.; Wu, Y.; Wan, X.; Li, J.; Tang, S.; Zhang, H.; and Zhuang, Y. 2025 · 2025
Closest in time.
EliGen: Entity-Level Controlled Image Generation with Regional Attention
Zhang, H.; Duan, Z.; Wang, X.; Chen, Y.; and Zhang, Y. 2025 · 2025
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C.; Yu, L.; Babu, A.; Tirumala, K.; Yasunaga, M.; Shamis, L.; Kahn, J.; Ma, X.; Zettlemoyer, L.; and Levy, O. 2025 · 2025
Closest in time.