Fetching the paper…
Reading the bibliography…
For multimodal LLMs, the synergy of visual comprehension (textual output) and generation (visual output) presents an ongoing challenge.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Earlier work this paper cites.
The book of why: the new science of cause and effect
Pearl, J. and Mackenzie, D · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Expressing visual relationships via language
Tan, H., Dernoncourt, F., Lin, Z., Bui, T., and Bansal, M · 2019
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D · 2020
Earlier work this paper cites.
Integrating multimodal information in large pretrained transformers
Rahman, W., Hasan, M. K., Lee, S., Zadeh, A., Mao, C., Morency, L.-P., and Hoque, E · 2020
Earlier work this paper cites.
Visual commonsense r-cnn
Wang, T., Huang, J., Zhang, H., and Sun, Q · 2020
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P., Qiu, L., Chen, J., Xia, T., Zhao, Y., Zhang, W., Yu, Z., Liang, X., and Zhu, S.-C · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Cited alongside, same era.
Learning by planning: Language-guided global image editing
Shi, J., Xu, N., Xu, Y., Bui, T., Dernoncourt, F., and Xu, C · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y · 2023
Later among the works it cites.
Planting a seed of vision in large language model
Ge, Y., Ge, Y., Zeng, Z., Wang, X., and Shan, Y · 2023
Later among the works it cites.
Unified language-vision pretraining with dynamic discrete visual tokenization
Jin, Y., Xu, K., Chen, L., Liao, C., Tan, J., Chen, B., Lei, C., Liu, A., Song, C., Lei, X., et al · 2023
Later among the works it cites.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R · 2023
Later among the works it cites.
Videopoet: A large language model for zero-shot video generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cao, M., Li, S., Li, J., Nie, L., and Zhang, M · 2022
Cited alongside, same era.
Laion coco: 600m synthetic captions from laion2b-en
Christoph, S., Andreas, K., Richard, V., Theo, C., and Romain, B · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Cited alongside, same era.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. Y · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Later among the works it cites.
Equivariant similarity for vision-language foundation models
Wang, T., Lin, K., Li, L., Lin, C.-C., Yang, Z., Zhang, H., Liu, Z., and Wang, L · 2023
Later among the works it cites.
Teal: Tokenize and embed all for multi-modal large language models
Yang, Z., Zhang, Y., Meng, F., and Zhou, J · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Later among the works it cites.
Magicbrush: A manually annotated dataset for instruction-guided image editing
Zhang, K., Mo, L., Chen, W., Sun, H., and Su, Y · 2023
Later among the works it cites.
Svit: Scaling up visual instruction tuning
Zhao, B., Wu, B., and Huang, T · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Worldgpt: Empowering llm as multimodal world model
Ge, Z., Huang, H., Zhou, M., Li, J., Wang, G., Tang, S., and Zhuang, Y · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y., Zhang, Y., Wang, C., Zhong, Z., Chen, Y., Chu, R., Liu, S., and Jia, J · 2024
Closest in time.
I3: Intent-introspective retrieval conditioned on instructions, 2024
Pan, K., Li, J., Wang, W., Fei, H., Song, H., Ji, W., Lin, J., Liu, X., Chua, T.-S., and Tang, S · 2024
Closest in time.
Anygpt: Unified multimodal llm with discrete sequence modeling
Zhan, J., Dai, J., Ye, J., Zhou, Y., Zhang, D., Liu, Z., Zhang, X., Yuan, R., Zhang, G., Li, L., et al · 2024
Closest in time.