Fetching the paper…
Reading the bibliography…
In this work, we aim to develop an MLLM that understands and solves questions by learning to create each intermediate step of the reasoning involved till the final answer.
Combining labeled and unlabeled data with co-training
Blum, A. and Mitchell, T · 1998
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Robust co-training
Sun, S. and Jin, F · 2011
Earlier work this paper cites.
Bayesian co-training
Yu, S., Krishnapuram, B., Rosales, R., and Rao, R. B · 2011
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Solving geometry problems: Combining text and diagram interpretation
Seo, M., Hajishirzi, H., Farhadi, A., Etzioni, O., and Malcolm, C · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., De Freitas, N., and Whiteson, S · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S. E., Michalski, V., Atkinson, A., Kádár, Á., Trischler, A., and Bengio, Y · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A., Seo, M., Schwenk, D., Choi, J., Farhadi, A., and Hajishirzi, H · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K., Price, B., Cohen, S., and Kanan, C · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J. J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D · 2018
Earlier work this paper cites.
Deep co-training for semi-supervised image recognition
Qiao, S., Shen, W., Zhang, Z., Wang, B., and Yuille, A · 2018
Earlier work this paper cites.
Maximum classifier discrepancy for unsupervised domain adaptation
Saito, K., Watanabe, K., Ushiku, Y., and Harada, T · 2018
Earlier work this paper cites.
Dec-mcts: Decentralized planning for multi-robot active perception
Best, G., Cliff, O. M., Patten, T., Mettu, R. R., and Fitch, R · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots
Methani, N., Ganguly, P., Khapra, M. M., and Kumar, P · 2020
Earlier work this paper cites.
Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Chen, J., Tang, J., Qin, J., Liang, X., Liu, L., Xing, E. P., and Lin, L · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C · 2021
Cited alongside, same era.
Mastering atari games with limited data
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y · 2021
Cited alongside, same era.
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Chen, J., Li, T., Qin, J., Lu, P., Lin, L., Chen, C., and Liang, X · 2022
Cited alongside, same era.
Genco: Generative co-training for generative adversarial networks with limited data
Cui, K., Huang, J., Luo, Z., Zhang, G., Zhan, F., and Lu, S · 2022
Cited alongside, same era.
Monte-carlo robot path planning
Dam, T., Chalvatzaki, G., Peters, J., and Pajarinen, J · 2022
Cited alongside, same era.
Discovering faster matrix multiplication algorithms with reinforcement learning
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., et al · 2024
Closest in time.
Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Dong, Y., Liu, Z., Sun, H.-L., Yang, J., Hu, W., Rao, Y., and Liu, Z · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., R Ruiz, F. J., Schrittwieser, J., Swirszcz, G., et al · 2022
Cited alongside, same era.
Hypertree proof search for neural theorem proving
Lample, G., Lacroix, T., Lachaux, M.-A., Rodriguez, A., Hayat, A., Lavril, T., Ebner, G., and Martinet, X · 2022
Cited alongside, same era.
Clevr-math: A dataset for compositional language, visual and mathematical reasoning
Lindström, A. D. and Abraham, S. S · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D. X., Tan, J. Q., Joty, S., and Hoque, E · 2022
Cited alongside, same era.
Infographicvqa
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C · 2022
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Huang, J. and Zhang, J · 2024
Closest in time.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Closest in time.
Building and better understanding vision-language models: insights and future directions
Laurençon, H., Marafioti, A., Sanh, V., and Tronchon, L · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Li, B., Zhang, K., Zhang, H., Guo, D., Zhang, R., Li, F., Zhang, Y., Liu, Z., and Li, C · 2024
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
Textcot: Zoom in for enhanced multimodal text-rich image understanding
Luan, B., Feng, H., Chen, H., Wang, Y., Zhou, W., and Li, H · 2024
Closest in time.
Improve mathematical reasoning in language models by automated process supervision
Luo, L., Liu, Y., Liu, R., Phatale, S., Lara, H., Li, Y., Shu, L., Zhu, Y., Meng, L., Sun, J., et al · 2024
Closest in time.
Compositional chain-of-thought prompting for large multimodal models
Mitra, C., Huang, B., Darrell, T., and Herzig, R · 2024
Closest in time.
Introducing openai o1, 2024
OpenAI · 2024
Closest in time.
Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
Shi, W., Hu, Z., Bin, Y., Liu, J., Yang, Y., Ng, S.-K., Bing, L., and Lee, R. K.-W · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Tong, S., Brown, E., Wu, P., Woo, S., Middepogu, M., Akula, S. C., Yang, J., Yang, S., Iyer, A., Pan, X., et al · 2024
Closest in time.
Vagadia, H., Chopra, M., Barnawal, A., Banerjee, T., Tuli, S., Chakraborty, S., and Paul, R · 2024
Closest in time.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding
Wu, Z., Chen, X., Pan, Z., Liu, X., Liu, W., Dai, D., Gao, H., Ma, Y., Wu, C., Wang, B., et al · 2024
Closest in time.
Monte carlo tree search boosts reasoning via iterative preference learning
Xie, Y., Goyal, A., Zheng, W., Kan, M.-Y., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Closest in time.
Llava-o1: Let vision language models reason step-by-step
Xu, G., Jin, P., Hao, L., Song, Y., Sun, L., and Yuan, L · 2024
Closest in time.
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al · 2024
Closest in time.
Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark
Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., et al · 2024
Closest in time.