Fetching the paper…
Reading the bibliography…
While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Gool, L. V., Gross, M. H., and Sorkine-Hornung, A · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X · 2018
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., Anderson, P., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with VQ-VAE-2
Razavi, A., van den Oord, A., and Vinyals, O · 2019
Earlier work this paper cites.
DM-GAN: dynamic memory generative adversarial networks for text-to-image synthesis
Zhu, M., Pan, P., Chen, W., and Yang, Y · 2019
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J · 2020
Earlier work this paper cites.
Diverse image generation via self-conditioned gans
Liu, S., Wang, T., Bau, D., Zhu, J., and Torralba, A · 2020
Earlier work this paper cites.
Are scene graphs good enough to improve image captioning?
Milewski, V. S. J., Moens, M., and Calixto, I · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
NVAE: A deep hierarchical variational autoencoder
Vahdat, A. and Kautz, J · 2020
Earlier work this paper cites.
Object relational graph with teacher-recommended learning for video captioning
Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., and Zha, Z · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Earlier work this paper cites.
A flow-based latent state generative model of neural population responses to natural images
Bashiri, M., Walker, E. Y., Lurz, K., Jagadish, A., Muhammad, T., Ding, Z., Ding, Z., Tolias, A. S., and Sinz, F. H · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., and Tang, J · 2021
Earlier work this paper cites.
Stochastic image-to-video synthesis using cinns
Dorkenwald, M., Milbich, T., Blattmann, A., Rombach, R., Derpanis, K. G., and Ommer, B · 2021
Earlier work this paper cites.
Automated audio captioning by fine-tuning BART with audioset tags
Gontier, F., Serizel, R., and Cerisara, C · 2021
Earlier work this paper cites.
Argmax flows and multinomial diffusion: Towards non-autoregressive language models
Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., and Welling, M · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W., Bolte, B., Tsai, Y. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T · 2021
Cited alongside, same era.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Pix2video: Video editing using image diffusion
Ceylan, D., Huang, C. P., and Mitra, N. J · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Closest in time.
Diffedit: Diffusion-based semantic image editing with mask guidance
Couairon, G., Verbeek, J., Schwenk, H., and Cord, M · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. C. H · 2023
Closest in time.
Cross-domain image captioning with discriminative finetuning
Dessì, R., Bevilacqua, M., Gualdoni, E., Rakotonirina, N. C., Franzon, F., and Baroni, M · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alayrac, J., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models, 2022
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V. Y., Huang, Y., Dai, A. M., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Cited alongside, same era.
Frido: Feature pyramid diffusion for complex scene image synthesis
Fan, W., Chen, Y., Chen, D., Cheng, Y., Yuan, L., and Wang, Y. F · 2022
Cited alongside, same era.
Training-free structured diffusion guidance for compositional text-to-image synthesis
Feng, W., He, X., Fu, T., Jampani, V., Akula, A. R., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Cited alongside, same era.
Long video generation with time-agnostic VQGAN and time-sensitive transformer
Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J., and Parikh, D · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Cited alongside, same era.
Dreamllm: Synergistic multimodal comprehension and creation
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., and Yi, L · 2023
Closest in time.
Layoutgpt: Compositional visual planning and generation with large language models
Feng, W., Zhu, W., Fu, T., Jampani, V., Akula, A. R., He, X., Basu, S., Wang, X. E., and Wang, W. Y · 2023
Closest in time.
Planting a SEED of vision in large language model
Ge, Y., Ge, Y., Zeng, Z., Wang, X., and Shan, Y · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I · 2023
Closest in time.
Text with knowledge graph augmented transformer for video captioning
Gu, X., Chen, G., Wang, Y., Zhang, L., Luo, T., and Wen, L · 2023
Closest in time.
Prompt-to-prompt image editing with cross-attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2023
Closest in time.
Dreampose: Fashion image-to-video synthesis via stable diffusion
Karras, J., Holynski, A., Wang, T., and Kemelmacher-Shlizerman, I · 2023
Closest in time.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R · 2023
Closest in time.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L · 2023
Closest in time.
Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A · 2023
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H. A., Khan, S. H., and Khan, F. S · 2023
Closest in time.
Mou, C., Wang, X., Xie, L., Zhang, J., Qi, Z., Shan, Y., and Qie, X · 2023
Closest in time.
SDXL: improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Closest in time.
Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D · 2023
Closest in time.
Generative pretraining in multimodality
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X · 2023
Closest in time.
Any-to-any generation via composable diffusion
Tang, Z., Yang, Z., Zhu, C., Zeng, M., and Bansal, M · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
P+: extended textual conditioning in text-to-image generation
Voynov, A., Chu, Q., Cohen-Or, D., and Aberman, K · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Closest in time.
mplug-2: A modularized multi-modal foundation model across text, image and video
Xu, H., Ye, Q., Yan, M., Shi, Y., Ye, J., Xu, Y., Li, C., Bi, B., Qian, Q., Wang, W., Xu, G., Zhang, J., Huang, S., Huang, F., and Zhou, J · 2023
Closest in time.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2023
Closest in time.
LAMM: language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., Shao, J., and Ouyang, W · 2023
Closest in time.
Unified language representation for question answering over text, tables, and images
Yu, B., Fu, C., Yu, H., Huang, F., and Li, Y · 2023
Closest in time.
Conzic: Controllable zero-shot image captioning by sampling-based polishing
Zeng, Z., Zhang, H., Lu, R., Wang, D., Chen, B., and Wang, Z · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Closest in time.