Fetching the paper…
Reading the bibliography…
As recent multi-modality large language models (MLLMs) have shown formidable proficiency on various complex tasks, there has been increasing attention on debating whether these models could eventually mirror human intelligence.
"general intelligence" objectively determined and measured
Spearman, C · 1961
Earlier work this paper cites.
Beyond IQ: A triarchic theory of human intelligence
Sternberg, R. J · 1985
Earlier work this paper cites.
Massive iq gains in 14 nations: What iq tests really measure
Flynn, J. R · 1987
Earlier work this paper cites.
Human Cognitive Abilities: A Survey of Factor-Analytic Studies
Carroll, J. B · 1993
Earlier work this paper cites.
A confirmatory factor analysis of the end-user computing satisfaction instrument
Doll, W. J., Xia, W., and Torkzadeh, G · 1994
Earlier work this paper cites.
Applications of structural equation modeling in marketing and consumer research: A review
Baumgartner, H. and Homburg, C · 1996
Earlier work this paper cites.
Raven progressive matrices
Raven, J · 2003
Earlier work this paper cites.
Internal and external factorial extensions to the cattell-horn-carroll (chc) theory of cognitive abilities: A review of factor analytic research since carroll’s seminal 1993 treatise
McGrew, K. S. and Evans, J. J · 2004
Earlier work this paper cites.
Essentials of Stanford-Binet intelligence scales (SB5) assessment , volume 39
Roid, G. H. and Barram, R. A · 2004
Earlier work this paper cites.
A proposal for the dartmouth summer research project on artificial intelligence, august 31, 1955
McCarthy, J., Minsky, M. L., Rochester, N., and Shannon, C. E · 2006
Earlier work this paper cites.
Chc theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research, 2009
McGrew, K. S · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
The stanford-binet intelligence scales , volume 654
Roid, G. H. and Pomplun, M · 2012
Earlier work this paper cites.
The cattell-horn-carroll model of intelligence
Schneider, W. J. and McGrew, K. S · 2012
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Neuropsychology: From theory to practice
Andrewes, D · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
From brain maps to cognitive ontologies: informatics and the search for mental structure
Poldrack, R. A. and Yarkoni, T · 2016
Earlier work this paper cites.
Essentials of WJ IV cognitive abilities assessment
Schrank, F. A., Decker, S. L., and Garruto, J. M · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A distributed brain network predicts general intelligence from resting-state human neuroimaging data
Dubois, J., Galdi, P., Paul, L. K., and Adolphs, R · 2018
Cited alongside, same era.
The cattell-horn-carroll theory of cognitive abilities
Schneider, W. J. and McGrew, K. S · 2018
Cited alongside, same era.
The woodcock–johnson iv
Schrank, F. A. and Wendling, B. J · 2018
Cited alongside, same era.
Exploratory factor analysis: A guide to best practice
Watkins, M. W · 2018
Cited alongside, same era.
Nocaps: Novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Cited alongside, same era.
Seed-bench: Benchmarking multimodal llms with generative comprehension
Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y · 2023
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Later among the works it cites.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., Meng, F., Huang, S., Qiao, Y., and Luo, P · 2023
Later among the works it cites.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang, Z., Li, L., Wang, J., Lin, K., Azarnasab, E., Ahmed, F., Liu, Z., Liu, C., Zeng, M., and Wang, L · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Cited alongside, same era.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Cited alongside, same era.
Concept formation and frontal lobe function: The search for a clinical frontal lobe test
Wang, P. L · 2019
Cited alongside, same era.
Beyond individual intelligence tests: application of cattell-horn-carroll theory
Caemmerer, J. M., Keith, T. Z., and Reynolds, M. R · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Cited alongside, same era.
Textcaps: a dataset for image captioning with reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Cited alongside, same era.
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Closest in time.
Cogbench: a large language model walks into a psychology lab
Coda-Forno, J., Binz, M., Wang, J. X., and Schulz, E · 2024
Closest in time.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2024
Closest in time.
Understanding social reasoning in language models with language models
Gandhi, K., Fränken, J.-P., Gerstenberg, T., and Goodman, N · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y., Zhang, Y., Wang, C., Zhong, Z., Chen, Y., Chu, R., Liu, S., and Jia, J · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., Sun, Y., et al · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., Bai, L., et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al · 2024
Closest in time.
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models
Zhang, W., Aljunied, M., Gao, C., Chia, Y. K., and Bing, L · 2024
Closest in time.