Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence.
Binary codes capable of correcting deletions, insertions, and reversals
Levenshtein, V. I. et al · 1966
Earlier work this paper cites.
Position bias in multiple-choice questions
Blunch, N. J · 1984
Earlier work this paper cites.
Thirteen ways to look at the correlation coefficient
Lee Rodgers, J. and Nicewander, W. A · 1988
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
A probabilistic interpretation of precision, recall and f-score, with implication for evaluation
Goutte, C. and Gaussier, E · 2005
Earlier work this paper cites.
Center-of-inattention: Position biases in decision-making
Raghubir, P. and Valenzuela, A · 2006
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Position bias estimation for unbiased learning to rank in personal search
Wang, X., Golbandi, N., Bendersky, M., Metzler, D., and Najork, M · 2018
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Evolution and impact of bias in human and machine learning algorithm interaction
Sun, W., Nasraoui, O., and Shafto, P · 2020
Earlier work this paper cites.
Infographicvqa
Mathew, M., Bagal, V., Tito, R. P., Karatzas, D., Valveny, E., and Jawahar, C · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Cited alongside, same era.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models
Wang, Z. J., Montoya, E., Munechika, D., Yang, H., Hoover, B., and Chau, D. H · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Openai models - gpt-4-vision
OpenAI · 2023
Later among the works it cites.
Are you spending too much money labeling data?, 2023
Prendki, J · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Verbosity bias in preference labeling by large language models
Saito, K., Wachi, A., Wataoka, K., and Akimoto, Y · 2023
Later among the works it cites.
Exploring ocr capabilities of gpt-4v (ision): A quantitative and in-depth evaluation
Shi, Y., Peng, D., Liao, W., Lin, Z., Chen, X., Liu, C., Zhang, Y., and Jin, L · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visit-bench: A benchmark for vision-language instruction following inspired by real-world use
Bitton, Y., Bansal, H., Hessel, J., Shao, R., Zhu, W., Awadalla, A., Gardner, J., Taori, R., and Schimdt, L · 2023
Cited alongside, same era.
Low-code llm: Visual programming over llms
Cai, Y., Mao, S., Wu, W., Wang, Z., Liang, Y., Ge, T., Wu, C., You, W., Song, T., Xia, Y., et al · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2023
Cited alongside, same era.
A survey of chain of thought reasoning: Advances, frontiers and future
Chu, Z., Chen, J., Chen, Q., Yu, W., He, T., Wang, H., Peng, W., Liu, M., Qin, B., and Liu, T · 2023
Cited alongside, same era.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., and Yao, H · 2023
Cited alongside, same era.
Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration
Deutsch, D., Foster, G., and Freitag, M · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models, 2023
GeminiTeam · 2023
Cited alongside, same era.
Aligning large multimodal models with factually augmented rlhf
Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., Yang, Y., et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Zhu, L., Wang, X., and Wang, X · 2023
Later among the works it cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Lee, S., Kim, S., Park, S. H., Kim, G., and Seo, M · 2024
Closest in time.
Leveraging large language models for nlg evaluation: A survey
Li, Z., Xu, X., Shen, T., Xu, C., Gu, J.-C., and Tao, C · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Closest in time.
Direct preference optimization of video large multimodal models from language model reward
Zhang, R., Gui, L., Sun, Z., Feng, Y., Xu, K., Zhang, Y., Fu, D., Li, C., Hauptmann, A., Bisk, Y., et al · 2024
Closest in time.