Fetching the paper…
Reading the bibliography…
We study the depth of grade-school math (GSM) problem-solving capabilities of LLMs.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
B. M. Lake and M. Baroni · 2018
Earlier work this paper cites.
Good-enough compositional data augmentation
J. Andreas · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
D. Hupkes, V. Dankers, M. Mul, and E. Bruni · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Reasoning chain based adversarial attack for multi-hop question answering
J. Ding, S. Wang, Q. Chen, and Z. Wei · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Earlier work this paper cites.
On the compositional generalization gap of in-context learning
A. Hosseini, A. Vani, D. Bahdanau, A. Sordoni, and A. C. Courville · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2023
Earlier work this paper cites.
PAL: program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Earlier work this paper cites.
Time travel in llms: Tracing data contamination in large language models
S. Golchin and M. Surdeanu · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
G. T. Google · 2023
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Z. Gou, Z. Shao, Y. Gong, Y. Yang, M. Huang, N. Duan, W. Chen, et al · 2023
Earlier work this paper cites.
Studying large language model generalization with influence functions
R. Grosse, J. Bae, C. Anil, N. Elhage, A. Tamkin, A. Tajdini, B. Steiner, D. Li, E. Durmus, E. Perez, et al · 2023
Earlier work this paper cites.
Decomposed prompting: A modular approach for solving complex tasks
T. Khot, H. Trivedi, M. Finlayson, Y. Fu, K. Richardson, P. Clark, and A. Sabharwal · 2023
Earlier work this paper cites.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis · 2023
Cited alongside, same era.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
N. Sardana and J. Frankle · 2023
Cited alongside, same era.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
N. Shapira, M. Levy, S. H. Alavi, X. Zhou, Y. Choi, Y. Goldberg, M. Sap, and V. Shwartz · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. H. Chi, N. Schärli, and D. Zhou · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, et al · 2023
Language models scale reliably with over-training and on downstream tasks
S. Y. Gadre, G. Smyrnis, V. Shankar, S. Gururangan, M. Wortsman, R. Shao, J. Mercat, A. Fang, J. Li, S. Keh, et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. T. Google · 2024
Closest in time.
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
T. He, D. Doshi, A. Das, and A. Gromov · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Are large language models really robust to word-level perturbations?
H. Wang, G. Ma, C. Yu, N. Gui, L. Zhang, Z. Huang, S. Ma, Y. Chang, S. Zhang, L. Shen, X. Wang, P. Zhao, and D. Tao · 2023
Cited alongside, same era.
Z. Wu, L. Qiu, A. Ross, E. Akyürek, B. Chen, B. Wang, N. Kim, J. Andreas, and Y. Kim · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. V. Le, and E. H. Chi · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl, A. Benhaim, M. Bilenko, J. Bjorck, S. Bubeck, Q. Cai, M. Cai, C. C. T. Mendes, W. Chen, V. Chaudhary, D. Chen, D. Chen, Y.-C. Chen, Y.-L. Chen, P. Chopra, X. Dai, A. D. Giorno, G. de Rosa, M. Dixon, R. Eldan, V. Fragoso, D. Iter, M. Gao, M. Gao, J. Gao, A. Garg, A. Goswami, S. Gunasekar, E. Haider, J. Hao, R. J. Hewett, J. Huynh, M. Javaheripi, X. Jin, P. Kauffmann, N. Karampatziakis, D. Kim, M. Khademi, L. Kurilenko, J. R. Lee, Y. T. Lee, Y. Li, Y. Li, C. Liang, L. Liden, C. Liu, M. Liu, W. Liu, E. Lin, Z. Lin, C. Luo, P. Madan, M. Mazzola, A. Mitra, H. Modi, A. Nguyen, B. Norick, B. Patra, D. Perez-Becker, T. Portet, R. Pryzant, H. Qin, M. Radmilac, C. Rosset, S. Roy, O. Ruwase, O. Saarikivi, A. Saied, A. Salim, M. Santacroce, S. Shah, N. Shang, H. Sharma, S. Shukla, X. Song, M. Tanaka, A. Tupini, X. Wang, L. Wang, C. Wang, Y. Wang, R. Ward, G. Wang, P. Witte, H. Wu, M. Wyatt, B. Xiao, C. Xu, J. Xu, W. Xu, S. Yadav, F. Yang, J. Yang, Z. Yang, Y. Yang, D. Yu, L. Yuan, C. Zhang, C. Zhang, J. Zhang, L. L. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, and X. Zhou · 2024
Cited alongside, same era.
On-policy distillation of language models: Learning from self-generated mistakes
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem · 2024
Cited alongside, same era.
M. Kazemi, H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut · 2024
Closest in time.
M. Levy, A. Jacoby, and Y. Goldberg · 2024
Closest in time.
M. Lewis and M. Mitchell · 2024
Closest in time.
Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers
Q. Li, L. Cui, X. Zhao, L. Kong, and W. Bi · 2024
Closest in time.
Unlocking tokens as data points for generalization bounds on larger language models
S. Lotfi, Y. Kuang, B. Amos, M. Goldblum, M. Finzi, and A. G. Wilson · 2024
Closest in time.
Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey
P. Mondorf and B. Plank · 2024
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Closest in time.
Ai-assisted generation of difficult math questions
V. Shah, D. Yu, K. Lyu, S. Park, N. R. Ke, M. Mozer, Y. Bengio, S. Arora, and A. Goyal · 2024
Closest in time.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
S. Srivastava, A. M. B, A. P. V, S. Menon, A. Sukumar, A. S. T, A. Philipose, S. Prince, and S. Thomas · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
G. Team M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Closest in time.
Efficient large language models: A survey
Z. Wan, X. Wang, C. Liu, S. Alam, Y. Zheng, J. Liu, Z. Qu, S. Yan, Y. Zhu, Q. Zhang, M. Chowdhury, and M. Zhang · 2024
Closest in time.
Benchmarking benchmark leakage in large language models
R. Xu, Z. Wang, R. Fan, and P. Liu · 2024
Closest in time.
On compositional generalization of transformer-based neural machine translation
Y. Yin, L. Fu, Y. Li, and Y. Zhang · 2024
Closest in time.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, S. Hendryx, R. Kaplan, M. Lunati, and S. Yue · 2024
Closest in time.