Fetching the paper…
Reading the bibliography…
Recent advancements in Multi-modal Large Language Models (MLLMs) have significantly improved their performance in tasks combining vision and language.
Youtube-8m: A large-scale video classification benchmark, 2016
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
Vqa: Visual question answering, 2016
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Florence: A new foundation model for computer vision, 2021
L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, C. Liu, M. Liu, Z. Liu, Y. Lu, Y. Shi, L. Wang, J. Wang, B. Xiao, Z. Xiao, J. Yang, M. Zeng, L. Zhou, and P. Zhang · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Z. Chen, Q. Zhou, Y. Shen, Y. Hong, H. Zhang, and C. Gan · 2023
Earlier work this paper cites.
Opencompass: A universal evaluation platform for foundation models
O. Contributors · 2023
Earlier work this paper cites.
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
D. Gao, L. Ji, L. Zhou, K. Q. Lin, J. Chen, Z. Fan, and M. Z. Shou · 2023
Earlier work this paper cites.
Towards mitigating LLM hallucination via self reflection
Z. Ji, T. Yu, Y. Xu, N. Lee, E. Ishii, and P. Fung · 2023
Earlier work this paper cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding, 2023
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
Vipergpt: Visual inference via python execution for reasoning
D. Surís, S. Menon, and C. Vondrick · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Image as a foreign language: Beit pretraining for vision and vision-language tasks
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, et al · 2023
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2023
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Llmeval: A preliminary study on how to evaluate large language models, 2023
Y. Zhang, M. Zhang, H. Yuan, S. Liu, Y. Shi, T. Gui, Q. Zhang, and X. Huang · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
Memory consolidation enables long-context video understanding, 2024
I. Balažević, Y. Shi, P. Papalampidi, R. Chaabouni, S. Koppula, and O. J. Hénaff · 2024
Closest in time.
Videoagent: A memory-augmented multimodal agent for video understanding, 2024
Y. Fan, X. Ma, R. Wu, Y. Du, J. Li, Z. Gao, and Q. Li · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Cited alongside, same era.
Large language models are better reasoners with self-verification, 2023
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, S. Liu, B. Sun, K. Liu, and J. Zhao · 2023
Cited alongside, same era.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models, 2023
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Cited alongside, same era.
A survey on multimodal large language models
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen · 2023
Cited alongside, same era.
Idealgpt: Iteratively decomposing vision and language reasoning via large language models
H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. A. Ayyubi, K.-W. Chang, and S.-F. Chang · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Cited alongside, same era.
Microsoft · 2024
Closest in time.
Azure openai service
Microsoft · 2024
Closest in time.
Reference video search - azure ai computer vision
Microsoft · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2024
Closest in time.
Critical thinking — wikipedia, the free encyclopedia, 2024
Wikipedia contributors · 2024
Closest in time.
A simple llm framework for long-range video question-answering, 2024
C. Zhang, T. Lu, M. M. Islam, Z. Wang, S. Yu, M. Bansal, and G. Bertasius · 2024
Closest in time.