Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel at various tasks, including solving math word problems (MWPs), but struggle with real-world problems containing irrelevant information.
Learning to automatically solve algebra word problems
Nate Kushman, Yoav Artzi, Luke Zettlemoyer, and Regina Barzilay · 2014
Earlier work this paper cites.
MAWPS: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi · 2016
Earlier work this paper cites.
Solving general arithmetic word problems
Subhro Roy and Dan Roth · 2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
Intent detection for code-mix utterances in task oriented dialogue systems
Pratik Jayarao and Aman Srivastava · 2018
Earlier work this paper cites.
Mapping to declarative knowledge for word problem solving
Subhro Roy and Dan Roth · 2018
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2020
Earlier work this paper cites.
Retraining distilbert for a voice shopping assistant by using universal dependencies, 2021
Pratik Jayarao and Arpit Sharma · 2021
Earlier work this paper cites.
Adversarial examples for evaluating math word problem solvers
Vivek Kumar, Rishabh Maheshwary, and Vikram Pudi · 2021
Earlier work this paper cites.
Deep kernels with probabilistic embeddings for small-data learning
Ankur Mallick, Chaitanya Dwivedi, Bhavya Kailkhura, Gauri Joshi, and T. Yong-Jin Han · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Multi stain graph fusion for multimodal integration in pathology
Chaitanya Dwivedi, Shima Nofallah, Maryam Pouryahya, Janani Iyer, Kenneth Leidal, Chuhan Chung, Timothy Watkins, Andrew Billin, Robert Myers, John Abel, and Ali Behrooz · 2022
Earlier work this paper cites.
Practice makes a solver perfect: Data augmentation for math word problem solvers
Vivek Kumar, Rishabh Maheshwary, and Vikram Pudi · 2022
Earlier work this paper cites.
LILA: A unified benchmark for mathematical reasoning
Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Yu-Chu Chang, Xu Wang, Jindong Wang, Yuanyi Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Weirong Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qian Yang, and Xingxu Xie · 2023
Cited alongside, same era.
Reasoning in large language models through symbolic math word problems
An llm can fool itself: A prompt-based adversarial attack
Xilie Xu, Keyi Kong, Ninghao Liu, Li zhen Cui, Di Wang, Jingfeng Zhang, and Mohan S. Kankanhalli · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges, 2024
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin · 2024
Closest in time.
Yi: Open foundation models by 01.ai, 2024
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
Au large, Apr 2024
Mistral AI · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vedant Gaur and Nikunj Saunshi · 2023
Cited alongside, same era.
“john is 50 years old, can his son be 65?” evaluating nlp models’ understanding of feasibility
Himanshu Gupta, Neeraj Varshney, Swaroop Mishra, Kuntal Kumar Pal, Saurabh Arjun Sawant, Kevin Scaria, Siddharth Goyal, and Chitta Baral · 2023
Cited alongside, same era.
MathPrompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Cited alongside, same era.
Maea: Multimodal attribution for embodied ai
Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Yonatan Bisk · 2023
Cited alongside, same era.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco J R Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, Alhussein Fawzi, Josh Grochow, Andrea Lodi, Jean-Baptiste Mouret, Talia Ringer, and Tao Yu · 2023
Cited alongside, same era.
Arb: Advanced reasoning benchmark for large language models
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, and Aran Komatsuzaki · 2023
Cited alongside, same era.
An independent evaluation of chatgpt on mathematical word problems (mwp), 2023
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Huai hsin Chi, Nathanael Scharli, and Denny Zhou · 2023
Cited alongside, same era.
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
Reka core, flash, and edge: A series of powerful multimodal language models, 2024
Aitor Ormazabal, Che Zheng, Cyprien de Masson d’Autume, Dani Yogatama, Deyu Fu, Donovan Ong, Eric Chen, Eugenie Lamprecht, Hai Pham, Isaac Ong, Kaloyan Aleksiev, Lei Li, Matthew Henderson, Max Bain, Mikel Artetxe, Nishant Relan, Piotr Padlewski, Qi Liu, Ren Chen, Samuel Phua, Yazheng Yang, Yi Tay, Yuqi Wang, Zhongkai Zhu, and Zhihui Xie · 2024
Closest in time.
Adversarial math word problem generation, 2024
Roy Xie, Chengxuan Huang, Junlin Wang, and Bhuwan Dhingra · 2024
Closest in time.
Mathattack: Attacking large language models towards math solving ability
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang · 2024
Closest in time.
Huemanity: Probing fine-grained visual perception in mllms
Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Nilay Pande · 2025
Closest in time.
Ai guide dog: Egocentric path prediction on smartphone
Aishwarya Jadhav, Jeffery Cao, Abhishree Shetty, Urvashi Kumar, Aditi Sharma, Ben Sukboontip, Jayant Tamarapalli, Jingyi Zhang, and Aniruddh Koul · 2025
Closest in time.
Countqa: How well do mllms count in the wild?
Jayant Sravan Tamarapalli, Rynaa Grover, Nilay Pande, and Sahiti Yerramilli · 2025
Closest in time.
Geochain: Multimodal chain-of-thought for geographic reasoning
Sahiti Yerramilli, Nilay Pande, Rynaa Grover, and Jayant Sravan Tamarapalli · 2025
Closest in time.