Fetching the paper…
Reading the bibliography…
Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Roscoe: A suite of metrics for scoring step-by-step reasoning
Golovneva, O., Chen, M., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Videollm: Modeling video sequence with large language models
Chen, G., Zheng, Y.-D., Wang, J., Xu, J., Huang, Y., Pan, J., Wang, Y., Wang, Y., Qiao, Y., Lu, T., et al · 2023
Earlier work this paper cites.
Guo, Z., Zhang, R., Zhu, X., Tang, Y., Ma, X., Han, J., Chen, K., Gao, P., Li, X., Li, H., et al · 2023
Earlier work this paper cites.
Videochat: Chat-centric video understanding
Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y · 2023
Earlier work this paper cites.
Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Earlier work this paper cites.
Lu, P., Bansal, H., Xia, T., Liu, J., yue Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Earlier work this paper cites.
GPT-4V(ision) system card, 2023
OpenAI · 2023
Earlier work this paper cites.
Receval: Evaluating reasoning chains via correctness and informativeness
Prasad, A., Saha, S., Zhou, X., and Bansal, M · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Earlier work this paper cites.
Pointllm: Empowering large language models to understand point clouds
Xu, R., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Cited alongside, same era.
Llava-grounding: Grounded visual chat with large multimodal models
Zhang, H., Li, H., Li, F., Ren, T., Zou, X., Liu, S., Huang, S., Gao, J., Zhang, L., Li, C., et al · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Cited alongside, same era.
Chimera: Improving generalist model with domain-specific experts
Peng, T., Li, M., Zhou, H., Xia, R., Zhang, R., Bai, L., Mao, S., Wang, B., He, C., Zhou, A., et al · 2024
Later among the works it cites.
Qwen2-vl
Qwen Team · 2024
Later among the works it cites.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Sprague, Z., Yin, F., Rodriguez, J. D., Jiang, D., Wadhwa, M., Singhal, P., Zhao, X., Ye, X., Mahowald, K., and Durrett, G · 2024
Later among the works it cites.
Qvq: To see the world with wisdom, December 2024
Team, Q · 2024
Later among the works it cites.
Llava-cot: Let vision language models reason step-by-step, 2024
Xu, G., Jin, P., Li, H., Song, Y., Sun, L., and Yuan, L · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., et al · 2024
Cited alongside, same era.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Gao, P., Zhang, R., Liu, C., Qiu, L., Huang, S., Lin, W., Zhao, S., Geng, S., Lin, Z., Jin, P., et al · 2024
Cited alongside, same era.
Hao, S., Gu, Y., Luo, H., Liu, T., Shao, X., Wang, X., Xie, S., Ma, H., Samavedhi, A., Gao, Q., et al · 2024
Cited alongside, same era.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Jia, Y., Liu, J., Chen, S., Gu, C., Wang, Z., Luo, L., Lee, L., Wang, P., Wang, Z., Zhang, R., et al · 2024
Cited alongside, same era.
Mmsearch: Benchmarking the potential of large models as multi-modal search engines
Jiang, D., Zhang, R., Guo, Z., Wu, Y., Lei, J., Qiu, P., Lu, P., Chen, Z., Song, G., Gao, P., et al · 2024
Cited alongside, same era.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Cited alongside, same era.
Introducing openai o1, 2024., 2024a
OpenAI · 2024
Cited alongside, same era.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z · 2024
Later among the works it cites.
Ying, K., Meng, F., Wang, J., Li, Z., Lin, H., Yang, Y., Zhang, H., Zhang, W., Lin, Y., Liu, S., et al · 2024
Later among the works it cites.
Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark, 2024
Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., Su, Y., Chen, W., and Neubig, G · 2024
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention
Zhang, R., Han, J., Liu, C., Zhou, A., Lu, P., Qiao, Y., Li, H., and Gao, P · 2024
Later among the works it cites.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Zhang, R., Jiang, D., Zhang, Y., Lin, H., Guo, Z., Qiu, P., Zhou, A., Lu, P., Chang, K.-W., Gao, P., et al · 2024
Later among the works it cites.
Virgo: A preliminary exploration on reproducing o1-like mllm
Du, Y., Liu, Z., Li, Y., Zhao, W. X., Huo, Y., Wang, B., Chen, W., Liu, Z., Wang, Z., and Wen, J.-R · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al · 2025
Closest in time.