Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have scaled up to unlock a wide range of complex reasoning tasks with the aid of various prompting methods.
Learning to solve arithmetic word problems with verb categorization
M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman · 2014
Earlier work this paper cites.
Parsing algebraic word problems into equations
R. Koncel-Kedziorski, H. Hajishirzi, A. Sabharwal, O. Etzioni, and S. D. Ang · 2015
Earlier work this paper cites.
Reasoning about quantities in natural language
S. Roy, T. Vieira, and D. Roth · 2015
Earlier work this paper cites.
MAWPS: A math word problem repository
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi · 2016
Earlier work this paper cites.
Solving general arithmetic word problems, 2016
S. Roy and D. Roth · 2016
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Earlier work this paper cites.
Semantically-aligned equation generation for solving and reasoning math word problems, 2019
T.-R. Chiang and Y.-N. Chen · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Are nlp models really able to solve simple math word problems?, 2021
A. Patel, S. Bhattamishra, and N. Goyal · 2021
Cited alongside, same era.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks, 2022
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Cited alongside, same era.
Solving math word problems by combining language models with symbolic solvers, 2023
J. He-Yueya, G. Poesia, R. E. Wang, and N. D. Goodman · 2023
Closest in time.
Decomposed prompting: A modular approach for solving complex tasks, 2023
T. Khot, H. Trivedi, M. Finlayson, Y. Fu, K. Richardson, P. Clark, and A. Sabharwal · 2023
Closest in time.
Large language models are zero-shot reasoners, 2023
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Is chatgpt a general-purpose natural language processing task solver?, 2023
C. Qin, A. Zhang, Z. Zhang, J. Chen, M. Yasunaga, and D. Yang · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring and narrowing the compositionality gap in language models, 2022
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications, 2022
R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, P. Srinivasan, L. Man, K. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le · 2022
Cited alongside, same era.
Teaching large language models to self-debug, 2023
X. Chen, M. Lin, N. Schärli, and D. Zhou · 2023
Cited alongside, same era.
Binding language models in symbolic languages, 2023
Z. Cheng, T. Xie, P. Shi, C. Li, R. Nadkarni, Y. Hu, C. Xiong, D. Radev, M. Ostendorf, L. Zettlemoyer, N. A. Smith, and T. Yu · 2023
Cited alongside, same era.
Complexity-based prompting for multi-step reasoning, 2023
Y. Fu, H. Peng, A. Sabharwal, P. Clark, and T. Khot · 2023
Cited alongside, same era.
Pal: Program-aided language models, 2023
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Compositional generalization via neural-symbolic stack machines, 2020a
X. Chen, C. Liang, A. W. Yu, D. Song, and D. Zhou
Cited in the paper.
Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension
X. Chen, C. Liang, A. W. Yu, D. Zhou, D. Song, and Q. V. Le
Cited in the paper.
Closest in time.
Self-consistency improves chain of thought reasoning in language models, 2023
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Closest in time.
Least-to-most prompting enables complex reasoning in large language models, 2023
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, and E. Chi · 2023
Closest in time.