Fetching the paper…
Reading the bibliography…
Multi-modality promises to unlock further uses for large language models.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Vqa: Visual question answering, 2016
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., and Radev, D · 2018
Earlier work this paper cites.
On the measure of intelligence
Chollet, F · 2019
Earlier work this paper cites.
A comprehensive survey of deep learning for image captioning
Hossain, M. Z., Sohel, F., Shiratuddin, M. F., and Laga, H · 2019
Cited alongside, same era.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Do, X. L., Tan, J. Q., Joty, S., and Hoque, E · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Cited alongside, same era.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2023
Cited alongside, same era.
The abstraction and reasoning corpus (arc), 2023
Chollet, F · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Closest in time.
Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models, 2023
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Closest in time.
OpenAI · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., and Tu, Z · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities
Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X
Cited in the paper.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H., Li, X., and Bing, L
Cited in the paper.
Multimodal chain-of-thought reasoning in language models
Zhang, Z., Zhang, A., Li, M., Zhao, H., Karypis, G., and Smola, A
Cited in the paper.
Closest in time.
Format5: Abstention and examples for conditional table formatting with natural language, 2023
Singh, M., Cambronero, J., Gulwani, S., Le, V., Negreanu, C., Nouri, E., Raza, M., and Verbruggen, G · 2023
Closest in time.
Next-gpt: Any-to-any multimodal llm
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S · 2023
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Closest in time.