Fetching the paper…
Reading the bibliography…
Unified generative models have shown remarkable performance in text and image generation.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Making the v in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Qwen-VL: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., et al · 2023
Earlier work this paper cites.
MobileVLM: A fast, reproducible and strong vision language assistant for mobile devices
Chu, X., Qiao, L., Lin, X., Xu, S., Yang, Y., Hu, Y., Wei, F., Zhang, X., Zhang, B., Wei, X., et al · 2023
Earlier work this paper cites.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Earlier work this paper cites.
DeepFloyd IF, 2023
DeepFloyd · 2023
Earlier work this paper cites.
Making llama see and draw with seed tokenizer
Ge, Y., Zhao, S., Zeng, Z., Ge, Y., Li, C., Wang, X., and Shan, Y · 2023
Earlier work this paper cites.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X · 2023
Cited alongside, same era.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Schindler, G., Hornung, R., Birodkar, V., Yan, J., Chiu, M.-C., et al · 2023
Cited alongside, same era.
Aesthetic predictor
LAION-AI · 2023
Cited alongside, same era.
Introducing IDEFICS: An open reproduction of state-of-the-art visual language model, 2023, 2023
Laurençon, H., van Strien, D., Bekman, S., Tronchon, L., Saulnier, L., Wang, T., Karamcheti, S., Singh, A., Pistilli, G., Jernite, Y., et al · 2023
Cited alongside, same era.
Self-planning code generation with large language models
Jiang, X., Dong, Y., Wang, L., Fang, Z., Shang, Q., Li, G., Jin, Z., and Jiao, W · 2024
Later among the works it cites.
Unified language-vision pretraining in llm with dynamic discrete visual tokenization
Jin, Y., Xu, K., Chen, L., Liao, C., Tan, J., Huang, Q., Bin, C., Song, C., ZHANG, D., Ou, W., et al · 2024
Later among the works it cites.
Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion
Ju, X., Liu, X., Wang, X., Bian, Y., Shan, Y., and Xu, Q · 2024
Later among the works it cites.
Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models
Liang, W., Yu, L., Luo, L., Iyer, S., Dong, N., Zhou, C., Ghosh, G., Lewis, M., Yih, W.-t., Zettlemoyer, L., et al · 2024
Later among the works it cites.
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G · 2023
Cited alongside, same era.
Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., and Li, H · 2023
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., et al · 2023
Cited alongside, same era.
4m-21: An any-to-any vision model for tens of tasks and modalities
Bachmann, R., Kar, O. F., Mizrahi, D., Garjani, A., Gao, M., Griffiths, D., Hu, J., Dehghan, A., and Zamir, A · 2024
Cited alongside, same era.
Later among the works it cites.
Ma, Y., Liu, X., Chen, X., Liu, W., Wu, C., Wu, Z., Pan, Z., Xie, Z., Zhang, H., Zhao, L., et al · 2024
Later among the works it cites.
Compositional chain-of-thought prompting for large multimodal models
Mitra, C., Huang, B., Darrell, T., and Herzig, R · 2024
Later among the works it cites.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Mou, C., Wang, X., Xie, L., Wu, Y., Zhang, J., Qi, Z., and Shan, Y · 2024
Later among the works it cites.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2024
Later among the works it cites.
Llamafusion: Adapting pretrained language models for multimodal generation
Shi, W., Han, X., Zhou, C., Liang, W., Lin, X. V., Zettlemoyer, L., and Yu, L · 2024
Later among the works it cites.
Metamorph: Multimodal understanding and generation via instruction tuning
Tong, S., Fan, D., Zhu, J., Xiong, Y., Chen, X., Sinha, K., Rabbat, M., LeCun, Y., Xie, S., and Liu, Z · 2024
Later among the works it cites.
Omnigen: Unified image generation
Xiao, S., Wang, Y., Zhou, J., Yuan, H., Xing, X., Yan, R., Wang, S., Huang, T., and Liu, Z · 2024
Later among the works it cites.
Llava-cot: Let vision language models reason step-by-step
Xu, G., Jin, P., Wu, Z., Li, H., Song, Y., Sun, L., and Yuan, L · 2024
Later among the works it cites.
LLaVA-Phi: Efficient multi-modal assistant with small language model
Zhu, Y., Zhu, M., Liu, N., Ou, Z., Mou, X., and Tang, J · 2024
Later among the works it cites.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al · 2025
Closest in time.
Devil is in the detail: Towards injecting fine details of image prompt in image generation via conflict-free guidance and stratified attention
Jo, K., Yun, J., and Choo, J · 2025
Closest in time.
Introducing gpt-5
OpenAI · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Zhao, Q., Lu, Y., Kim, M. J., Fu, Z., Zhang, Z., Wu, Y., Li, Z., Ma, Q., Han, S., Finn, C., et al · 2025
Closest in time.