Fetching the paper…
Reading the bibliography…
With the continuous advancement of large language models (LLMs), it is essential to create new benchmarks to effectively evaluate their expanding capabilities and identify areas for improvement.
On random graphs
P. Erdős and A. Rényi · 1959
Earlier work this paper cites.
Stochastic blockmodels: First steps
P. W. Holland, K. B. Laskey, and S. Leinhardt · 1983
Earlier work this paper cites.
Emergence of scaling in random networks
A.-L. Barabási and R. Albert · 1999
Earlier work this paper cites.
Statistical mechanics of complex networks
R. Albert and A.-L. Barabási · 2002
Earlier work this paper cites.
Exploring network structure, dynamics, and function using networkx
A. Hagberg, P. Swart, and D. S Chult · 2008
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
N. F. Rajani, B. McCann, C. Xiong, and R. Socher · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Visually grounded reasoning across languages and cultures
F. Liu, E. Bugliarello, E. M. Ponti, S. Reddy, N. Collier, and D. Elliott · 2021
Cited alongside, same era.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu · 2021
Cited alongside, same era.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
O. Tafjord, B. Dalvi, and P. Clark · 2021
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2023
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multimodal few-shot learning with frozen language models
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Cited alongside, same era.
Clevr-math: A dataset for compositional language, visual and mathematical reasoning
A. D. Lindström and S. S. Abraham · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
Language models are multilingual chain-of-thought reasoners
F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, et al · 2022
Cited alongside, same era.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2023
Later among the works it cites.
The curious case of nonverbal abstract reasoning with multi-modal large language models
K. Ahrabian, Z. Sourati, K. Sun, J. Zhang, Y. Jiang, F. Morstatter, and J. Pujara · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
A. Anthropic · 2024
Closest in time.
PaliGemma: A versatile 3B VLM for transfer, 2024
L. Beyer*, A. Steiner*, A. Susano Pinto*, A. Kolesnikov*, X. Wang*, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, A. Gritsenko, X. Chen, S. Koppula, A. Grycner, M. Bauer, M. Bošnjak, F. Liu, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai* · 2024
Closest in time.
Nphardeval4v: A dynamic reasoning benchmark of multimodal large language models
L. Fan, W. Hua, X. Li, K. Zhu, M. Jin, L. Li, H. Ling, J. Chi, J. Wang, X. Ma, et al · 2024
Closest in time.
Talk like a graph: Encoding graphs for large language models
B. Fatemi, J. Halcrow, and B. Perozzi · 2024
Closest in time.
Blink: Multimodal large language models can see but not perceive
X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna · 2024
Closest in time.
Language is not all you need: Aligning perception with language models
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, et al · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Testing the general deductive reasoning capacity of large language models using ood examples
A. Saparov, R. Y. Pang, V. Padmakumar, N. Joshi, M. Kazemi, N. Kim, and H. He · 2024
Closest in time.
A corpus of natural language for visual reasoning
A. Suhr, M. Lewis, J. Yeh, and Y. Artzi · 2034
Closest in time.