Fetching the paper…
Reading the bibliography…
Over the past years, advances in artificial intelligence (AI) have demonstrated how AI can solve many perception and generation tasks, such as image classification and text writing, yet reasoning remains a challenge.
Socialiqa: Commonsense reasoning about social interactions, 2019
Sap, M., Rashkin, H., Chen, D., LeBras, R. and Choi, Y · 1904
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models, 2019
Saxton, D., Grefenstette, E., Hill, F. and Kohli, P · 1904
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. and Choi, Y · 1905
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R.L., Gao, J. and Choi, Y · 1911
Earlier work this paper cites.
On the measure of intelligence, 2019
Chollet, F · 1911
Earlier work this paper cites.
Meshed-memory transformer for image captioning, 2020
Cornia, M., Stefanini, M., Baraldi, L. and Cucchiara, R · 1912
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T.B. et al · 2005
Earlier work this paper cites.
Text modular networks: Learning to decompose tasks in the language of existing models, 2021
Khot, T., Khashabi, D., Richardson, K., Clark, P. and Sabharwal, A · 2009
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E. and Morgenstern, L · 2012
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
He, K., Zhang, X., Ren, S. and Sun, J · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C.L. and Girshick, R · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C. and Tafjord, O · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S. and Soricut, R · 2018
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs, 2019
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H. and Artzi, Y · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Talmor, A., Herzig, J., Lourie, N. and Berant, J · 2019
Earlier work this paper cites.
Unifying vision-and-language tasks via text generation, 2021
Cho, J., Lei, J., Tan, H. and Bansal, M · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K. et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D. and Steinhardt, J · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models, 2021
Nye, M. et al · 2021
Cited alongside, same era.
A plug-and-play method for controlled text generation
Pascual, D., Egressy, B., Meister, C., Cotterell, R. and Wattenhofer, R · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Radford, A. et al · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning, 2022
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Abdin, M. et al · 2024
Later among the works it cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
Glazer, E. et al · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Grattafiori, A. et al · 2024
Later among the works it cites.
Grin: Gradient-informed moe, 2024
Liu, L. et al · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alayrac, J.B. et al · 2022
Cited alongside, same era.
On the opportunities and risks of foundation models, 2022
Bommasani, R. et al · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners, 2022
Wei, J., Bosma, M., Zhao, V.Y., Guu, K., Yu, A.W., Lester, B., Du, N., Dai, A.M. and Le, Q.V · 2022
Cited alongside, same era.
Making large multimodal models understand arbitrary visual prompts, 2023
Cai, M., Liu, H., Mustikovela, S.K., Meyer, G.P., Chai, Y., Park, D. and Lee, Y.J · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model, 2023
Driess, D. et al · 2023
Cited alongside, same era.
Ultralytics yolov8, 2023
Jocher, G., Chaurasia, A. and Qiu, J · 2023
Cited alongside, same era.
The defeat of the winograd schema challenge, 2023
Kocijan, V., Davis, E., Lukasiewicz, T., Marcus, G. and Morgenstern, L · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners, 2023
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y. and Iwasawa, Y · 2023
Cited alongside, same era.
Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S. and Farajtabar, M · 2024
Later among the works it cites.
Chatgpt: Language model by openai, 2023
OpenAI · 2024
Later among the works it cites.
OpenAI et al · 2024
Later among the works it cites.
Breaking reCAPTCHAv2
Plesner, A., Vontobel, T. and Wattenhofer, R · 2024
Later among the works it cites.
Mutual reasoning makes smaller llms stronger problem-solvers, 2024
Qi, Z., Ma, M., Xu, J., Zhang, L.L., Yang, F. and Yang, M · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K. and Kumar, A · 2024
Later among the works it cites.
Yue, X. et al · 2024
Later among the works it cites.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking, 2025
Guan, X., Zhang, L.L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F. and Yang, M · 2025
Closest in time.
Idena Blockchain explorer, 2025a
Idena · 2025
Closest in time.
What is a flip?, 2025b
Idena · 2025
Closest in time.
Idena Whitepaper, 2025c
Idena · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen et al · 2025
Closest in time.