Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., and Van Den Hengel, A · 2018
Earlier work this paper cites.
Geometric multimodal representation learning
Ektefaie, Y., Dasoulas, G., Noori, A., Farhat, M., and Zitnik, M · 2022
Earlier work this paper cites.
Vision-and-language navigation: A survey of tasks, methods, and future directions
Gu, J., Stefani, E., Wu, Q., Thomason, J., and Wang, X. E · 2022
Earlier work this paper cites.
Self-instruct: Aligning language model with self generated instructions, 2022
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Llm-in-the-loop: Leveraging large language model for thematic analysis
Dai, S.-C., Xiong, A., and Ku, L.-W · 2023
Cited alongside, same era.
Mathprompter: Mathematical reasoning using large language models
Imani, S., Du, L., and Shrivastava, H · 2023
Cited alongside, same era.
A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging
Li, Y., Liu, Y., Wang, Z., Liang, X., Liu, L., Wang, L., Cui, L., Tu, Z., Wang, L., and Zhou, L · 2023
Cited alongside, same era.
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
OpenAI · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration
Wake, N., Kanehira, A., Sasabuchi, K., Takamatsu, J., and Ikeuchi, K · 2023
Later among the works it cites.
On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S
Cited in the paper.
Lmeye: An interactive perception network for large language models
Li, Y., Hu, B., Chen, X., Ma, L., and Zhang, M
Cited in the paper.
A comprehensive evaluation of gpt-4v on knowledge-intensive visual question answering
Li, Y., Wang, L., Hu, B., Chen, X., Zhong, W., Lyu, C., and Zhang, M
Cited in the paper.
Can language models solve graph problems in natural language?
Wang, H., Feng, S., He, T., Tan, Z., Han, X., and Tsvetkov, Y
Cited in the paper.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al
Cited in the paper.
Wen, L., Yang, X., Fu, D., Wang, X., Cai, P., Li, X., Ma, T., Li, Y., Xu, L., Shang, D., et al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Webvln: Vision-and-language navigation on websites
Chen, Q., Pitawela, D., Zhao, C., Zhou, G., Chen, H.-T., and Wu, Q · 2024
Closest in time.