Fetching the paper…
Reading the bibliography…
Recent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable AI models to outperform human abilities in varied tasks that demand higher-order cognitive skills.
The selection of upper and lower groups for the validation of test items
Truman Lee Kelley · 1939
Earlier work this paper cites.
The recognition problem. tech. rep
MM Bongard · 1968
Earlier work this paper cites.
Statistical Theories of Mental Test Scores
Frederic M. Lord and Melvin R. Novick · 1968
Earlier work this paper cites.
Society of mind
Marvin Minsky · 1988
Earlier work this paper cites.
Essentials of Educational Measurement
Robert L. Ebel and David A. Frisbie · 1991
Earlier work this paper cites.
The role of deliberate practice in the acquisition of expert performance
K Anders Ericsson, Ralf T Krampe, and Clemens Tesch-Römer · 1993
Earlier work this paper cites.
The power of feedback
John Hattie and Helen Timperley · 2007
Earlier work this paper cites.
A collection of definitions of intelligence
Shane Legg, Marcus Hutter, et al · 2007
Earlier work this paper cites.
Classroom Assessment: What Teachers Need to Know
W. James Popham · 2010
Earlier work this paper cites.
Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning
Elizabeth L Bjork, Robert A Bjork, et al · 2011
Earlier work this paper cites.
High achiever! Always a high achiever?: A comparison of student achievements on mathematical tests with different aims and goals
Elisabet Mellroth · 2014
Earlier work this paper cites.
Problem solving competency and the mathematical kangaroo
Elisabet Mellroth · 2015
Earlier work this paper cites.
The measure of all minds: evaluating natural and artificial intelligence
José Hernández-Orallo · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Girls’ performance in the kangaroo contest
Mark Applebaum and Roza Leikin · 2019
Cited alongside, same era.
On the measure of intelligence
François Chollet · 2019
Cited alongside, same era.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach · 2019
Cited alongside, same era.
How to create and solve: Analysis of items from the mathematical kangaroo from two perspectives
Lukas Andritsch, Evita Hauke, and Jakob Kelz · 2020
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2023
Later among the works it cites.
https://mathkangaroo.org/mks/, 2012–2024
Math Kangaroo USA, NFP Inc · 2024
Closest in time.
Claude-3 Opus
Anthropic · 2024
Closest in time.
InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al · 2024
Closest in time.
Gemini pro
Google DeepMind · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scale-localized abstract reasoning
Yaniv Benny, Niv Pekar, and Lior Wolf · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Visit-bench: A benchmark for vision-language instruction following inspired by real-world use
Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, and Ludwig Schimdt · 2023
Cited alongside, same era.
Are deep neural networks smarter than second graders?
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith, and Joshua B. Tenenbaum · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, et al · 2024
Closest in time.
Meng Fang, Xiangpeng Wan, Fei Lu, Fei Xing, and Kai Zou · 2024
Closest in time.
MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2024
Closest in time.
GSM-Symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Closest in time.
xgen-mm-phi3-mini-instruct model card, May 2024
Salesforce AI Research · 2024
Closest in time.
Measuring multimodal mathematical reasoning with math-vision dataset, 2024
Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Mingjie Zhan, and Hongsheng Li · 2024
Closest in time.
Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark
Minxuan Zhou, Hao Liang, Tianpeng Li, Zhiyu Wu, Mingan Lin, Linzhuang Sun, Yaqi Zhou, Yan Zhang, Xiaoqin Huang, Yicong Chen, et al · 2024
Closest in time.