Fetching the paper…
Reading the bibliography…
Large language models (LLMs) can solve an increasing number of complex reasoning tasks while making surprising mistakes in basic numerical understanding and processing (such as 9.11 > 9.9).
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Earlier work this paper cites.
Numeracy for language models: Evaluating and improving their ability to predict numbers
Georgios Spithourakis and Sebastian Riedel · 2018
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Exploring numeracy in word embeddings
Aakanksha Naik, Abhilasha Ravichander, Carolyn Rose, and Eduard Hovy · 2019
Earlier work this paper cites.
EQUATE: A benchmark evaluation framework for quantitative reasoning in natural language inference
Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2019
Earlier work this paper cites.
Do NLP models know numbers? probing numeracy in embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Injecting numerical reasoning skills into language models
Mor Geva, Ankit Gupta, and Jonathan Berant · 2020
Earlier work this paper cites.
Probing for multilingual numerical understanding in transformer-based language models
Devin Johnson, Denise Mak, Andrew Barker, and Lexi Loessberg-Zahl · 2020
Earlier work this paper cites.
Birds have four legs?! numersense: Probing numerical commonsense knowledge of pre-trained language models
Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren · 2020
Earlier work this paper cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2020
Earlier work this paper cites.
On long-tailed phenomena in neural machine translation
Vikas Raunak, Siddharth Dalmia, Vivek Gupta, and Florian Metze · 2020
Earlier work this paper cites.
Finqa: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R Routledge, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering
Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang · 2022
Cited alongside, same era.
Transformer language models without positional encodings still learn positional information
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh · 2022
Cited alongside, same era.
Low-dimension-to-high-dimension generalization and its implications for length generalization, 2024
Yang Chen, Yitao Liang, and Zhouchen Lin · 2024
Closest in time.
Position coupling: Leveraging task structure for improved length generalization of transformers
Hanseul Cho, Jaeyoung Cha, Pranjal Awasthi, Srinadh Bhojanapalli, Anupam Gupta, and Chulhee Yun · 2024
Closest in time.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Case-based or rule-based: how do transformers do the math?
Yi Hu, Xiaojuan Tang, Haotong Yang, and Muhan Zhang · 2024
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Exploring the numerical reasoning capabilities of language models: A comprehensive analysis on tabular data
Mubashara Akhtar, Abhilash Shankarampeta, Vivek Gupta, Arpit Patil, Oana Cocarascu, and Elena Simperl · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2023
Cited alongside, same era.
A survey of deep learning for mathematical reasoning
Pan Lu, Liang Qiu, Wenhao Yu, Sean Welleck, and Kai-Wei Chang · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Cited alongside, same era.
Closest in time.
Llms can find mathematical reasoning mistakes by pedagogical chain-of-thought
Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li · 2024
Closest in time.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.
Evaluating mathematical reasoning of large language models: A focus on error identification and correction
Xiaoyuan Li, Wenjie Wang, Moxin Li, Junrong Guo, Yang Zhang, and Fuli Feng · 2024
Closest in time.
MistralAI · 2024
Closest in time.
Alibaba Group Qwen Team · 2024
Closest in time.
Improving self consistency in LLMs through probabilistic tokenization
Ashutosh Sathe, Divyanshu Aggarwal, and Sunayana Sitaram · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Aaditya K Singh and DJ Strouse · 2024
Closest in time.
Mathscale: scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei · 2024
Closest in time.
Conveyor: Efficient tool-aware llm serving with tool partial execution
Yechen Xu, Xinhao Kong, Tingjun Chen, and Danyang Zhuo · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Yongchao Zhou, Uri Alon, Xinyun Chen, Xuezhi Wang, Rishabh Agarwal, and Denny Zhou · 2024
Closest in time.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong · 2025
Closest in time.