Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have recently demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time ("few-shot prompting").
An investigation of procedure and variable names as beacons during program comprehension
Gellenbeck, E. M. and Cook, C. R · 1991
Earlier work this paper cites.
The effects of comments and identifier names on program comprehensibility: an experimental investigation
Takang, A. A., Grubb, P. A., and Macredie, R. D · 1996
Earlier work this paper cites.
Mawps: A math word problem repository
Koncel-Kedziorski, R., Roy, S., Amini, A., Kushman, N., and Hajishirzi, H · 2016
Earlier work this paper cites.
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Deep learning: A critical appraisal
Marcus, G · 2018
Earlier work this paper cites.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H · 2019
Earlier work this paper cites.
Giving bert a calculator: Finding operations and arguments with reading comprehension
Andor, D., He, L., Lee, K., and Pitler, E · 2019
Earlier work this paper cites.
Neural module networks for reasoning over text
Gupta, N., Lin, K., Roth, D., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Just add functions: A neural-symbolic language model
Demeter, D. and Downey, D · 2020
Earlier work this paper cites.
Neurosymbolic ai: the 3rd wave
Garcez, A. d. and Lamb, L. C · 2020
Earlier work this paper cites.
The next decade in ai: four steps towards robust artificial intelligence
Marcus, G · 2020
Earlier work this paper cites.
A diverse corpus for evaluating and developing English math word problem solvers
Miao, S.-y., Liang, C.-C., and Su, K.-Y · 2020
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
Gehrmann, S., Adewumi, T., Aggarwal, K., Ammanamanchi, P. S., Anuoluwapo, A., Bosselut, A., Chandu, K. R., Clinciu, M., Das, D., Dhole, K. D., Du, W., Durmus, E., Dušek, O., Emezue, C., Gangal, V., Garbacea, C., Hashimoto, T., Hou, Y., Jernite, Y., Jhamtani, H., Ji, Y., Jolly, S., Kale, M., Kumar, D., Ladhak, F., Madaan, A., Maddela, M., Mahajan, K., Mahamood, S., Majumder, B. P., Martins, P. H., McMillan-Major, A., Mille, S., van Miltenburg, E., Nadeem, M., Narayan, S., Nikolaev, V., Niyongabo, R. A., Osei, S., Parikh, A., Perez-Beltrachini, L., Rao, N. R., Raunak, V., Rodriguez, J. D., Santhanam, S., Sedoc, J., Sellam, T., Shaikh, S., Shimorina, A., Cabezudo, M. A. S., Strobelt, H., Subramani, N., Xu, W., Yang, D., Yerukola, A., and Zhou, J · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the MATH dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Binding language models in symbolic languages
Cheng, Z., Xie, T., Shi, P., Li, C., Nadkarni, R., Hu, Y., Xiong, C., Radev, D., Ostendorf, M., Zettlemoyer, L., Smith, N. A., and Yu, T · 2022
Closest in time.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Closest in time.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2021
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks
Nogueira, R., Jiang, Z., and Lin, J · 2021
Cited alongside, same era.
Show your Work: Scratchpads for Intermediate Computation with Language Models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Cited alongside, same era.
Are NLP Models Really Able to Solve Simple Math Word Problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Cited alongside, same era.
A Recipe for Arbitrary Text Style Transfer with Large Language Models
Reif, E., Ippolito, D., Yuan, A., Coenen, A., Callison-Burch, C., and Wei, J · 2021
Cited alongside, same era.
Multitask Prompted Training Enables Zero-Shot Task Generalization , 2021
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Biderman, S., Gao, L., Bers, T., Wolf, T., and Rush, A. M · 2021
Cited alongside, same era.
Few-shot semantic parsing with language models trained on code
Shin, R. and Van Durme, B · 2021
Cited alongside, same era.
Constrained language models yield few-shot semantic parsers
Shin, R., Lin, C. H., Thomson, S., Chen, C., Roy, S., Platanios, E. A., Pauls, A., Klein, D., Eisner, J., and Van Durme, B · 2021
Cited alongside, same era.
Finetuned Language Models are Zero-shot Learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Madaan, A. and Yazdanbakhsh, A · 2022
Closest in time.
Language models of code are few-shot commonsense learners
Madaan, A., Zhou, S., Alon, U., Yang, Y., and Neubig, G · 2022
Closest in time.
Lila: A unified benchmark for mathematical reasoning
Mishra, S., Finlayson, M., Lu, P., Tang, L., Welleck, S., Baral, C., Rajpurohit, T., Tafjord, O., Sabharwal, A., Clark, P., and Kalyan, A · 2022
Closest in time.
Reasoning like program executors
Pi, X., Liu, Q., Chen, B., Ziyadi, M., Lin, Z., Gao, Y., Fu, Q., Lou, J.-G., and Chen, W · 2022
Closest in time.
Limitations of language models in arithmetic and symbolic induction
Qian, J., Wang, H., Li, Z., Li, S., and Yan, X · 2022
Closest in time.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Scharli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E., Zhou, D., and Wei, J · 2022
Closest in time.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Closest in time.
Autoformalization with Large Language Models
Wu, Y., Jiang, A. Q., Li, W., Rabe, M. N., Staats, C., Jamnik, M., and Szegedy, C · 2022
Closest in time.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Closest in time.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Bousquet, O., Le, Q., and Chi, E · 2022
Closest in time.