Fetching the paper…
Reading the bibliography…
We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC).
Effects of noise letters upon the identification of a target letter in a nonsearch task
Barbara A. Eriksen and Charles W. Eriksen. 1974 · 1974
Earlier work this paper cites.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2021 · 2012
Earlier work this paper cites.
ProtoQA: A question answering dataset for prototypical common-sense reasoning
Michael Boratko, Xiang Li, Tim O’Gorman, Rajarshi Das, Dan Le, and Andrew McCallum. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Eric Zelikman, Ran Tao, Tatsunori Hashimoto, John Schulman, Xinyun Chen, C Lawrence Zitnick, Christopher D Manning, Daniel Saunders, Adam Santoro, et al. 2022 · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, H. Francis Song, Noah Y. Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. 2022 · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022 · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Xuezhi Zhou, Nathanael Schärli, Luheng Hou, Jason Wei, Swaroop Mishra, Xinyun Huang, Quoc Le, Hyung Won Chung, Hieu Pham, Barret Zoph, et al. 2022 · 2022
Earlier work this paper cites.
Art: Automatic multi-step reasoning and tool-use for large language models
Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. 2023a · 2023
Earlier work this paper cites.
Language models can improve event prediction by few-shot abductive reasoning
Xiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou, James Y. Zhang, Jun Zhou, Chenhao Tan, and Hongyuan Mei. 2023b · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Qa-lora: Quantization-aware low-rank adaptation of large language models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, and Qi Tian. 2023 · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023b · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023c · 2023
Cited alongside, same era.
Cutting through the noise: Boosting llm performance on math word problems
Ujjwala Anantheswaran, Himanshu Gupta, Kevin Scaria, Shreyas Verma, Chitta Baral, and Swaroop Mishra. 2024 · 2024
Cited alongside, same era.
The reversal curse: Llms trained on "a is b" fail to learn "b is a"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2024 · 2024
Cited alongside, same era.
A systematic analysis of large language models as soft reasoners: The case of syllogistic inferences
Rag-instruct: Boosting llms with diverse retrieval-augmented instructions
Wanlong Liu, Junying Chen, Ke Ji, Li Zhou, Wenyu Chen, and Benyou Wang. 2024 · 2024
Later among the works it cites.
How much do prompting methods help llms on quantitative reasoning with irrelevant information?
Seok Hwan Song and Wallapak Tavanapong. 2024 · 2024
Later among the works it cites.
How easily do irrelevant inputs skew the responses of large language models?
Siye Wu, Jian Xie, Jiangjie Chen, Tinghui Zhu, Kai Zhang, and Yanghua Xiao. 2024 · 2024
Later among the works it cites.
Pride and prejudice: Llm amplifies self-bias in self-refinement
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Yang Wang. 2024 · 2024
Later among the works it cites.
Processbench: Identifying process errors in mathematical reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leonardo Bertolazzi, Albert Gatt, and Raffaella Bernardi. 2024 · 2024
Cited alongside, same era.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024 · 2024
Cited alongside, same era.
MC-indexing: Effective long document retrieval via multi-view content-aware indexing
Kuicai Dong, Derrick Goh Xin Deik, Yi Quan Lee, Hao Zhang, Xiangyang Li, Cong Zhang, and Yong Liu. 2024 · 2024
Cited alongside, same era.
LLM reasoners: New evaluation, library, and analysis of step-by-step reasoning with large language models
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, Zhen Wang, and Zhiting Hu. 2024 · 2024
Cited alongside, same era.
V-star: Training verifiers for self-taught reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron C. Courville, Alessandro Sordoni, and Rishabh Agarwal. 2024 · 2024
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024 · 2024
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, and Aleksandra Faust. 2024 · 2024
Cited alongside, same era.
Llms for relational reasoning: How far are we?
Zhiming Li, Yushi Cao, Xiufeng Xu, Junzhe Jiang, Xu Liu, Yon Shin Teo, Shang-Wei Lin, and Yang Liu. 2024 · 2024
Cited alongside, same era.
Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2024 · 2024
Later among the works it cites.
Lora-par: A flexible dual-system lora partitioning approach to efficient llm fine-tuning
Yining Huang, Bin Li, Keke Tang, and Meilian Chen. 2025 · 2025
Closest in time.
Vc search: Bridging the gap between well-defined and ill-defined problems in mathematical reasoning
Shi-Yu Tian, Zhi Zhou, Kun-Yang Yu, Ming Yang, Lin-Han Jia, Lan-Zhe Guo, and Yu-Feng Li. 2025 · 2025
Closest in time.
A survey of link prediction in n-ary knowledge graphs
Jiyao Wei, Saiping Guan, Da Li, Xiaolong Jin, Jiafeng Guo, and Xueqi Cheng. 2025 · 2025
Closest in time.
Mitigating posterior salience attenuation in long-context LLMs with positional contrastive decoding
Zikai Xiao, Ziyang Wang, Wen Ma, Yan Zhang, Wei Shen, WangYan WangYan, Luqi Gong, and Zuozhu Liu. 2025 · 2025
Closest in time.
Feng Xiong, Hongling Xu, Yifei Wang, Runxi Cheng, Yong Wang, and Xiangxiang Chu. 2025 · 2025
Closest in time.
Utmath: Math evaluation with unit test via reasoning-to-coding thoughts
Bo Yang, Qingping Yang, Yingwei Ma, and Runtao Liu. 2025 · 2025
Closest in time.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process
Tian Ye, Zicheng Xu, Yuanzhi Li, and Zeyuan Allen-Zhu. 2025 · 2025
Closest in time.
Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian, and Beidi Chen. 2025 · 2025
Closest in time.