Fetching the paper…
Reading the bibliography…
While large language models (LLMs) such as ChatGPT and PaLM have demonstrated remarkable performance in various language understanding and generation tasks, their capabilities in complex reasoning and intricate knowledge utilization still fall short of human-level proficiency.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Amini, A.; Gabriel, S.; Lin, P.; Koncel-Kedziorski, R.; Choi, Y.; and Hajishirzi, H. 2019 · 1905
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T.; Fisch, A.; and Chen, D. 2020 · 2012
Earlier work this paper cites.
GPT2: Empirical slant delay model for radio space geodetic techniques
Lagler, K.; Schindelegger, M.; Böhm, J.; Krásná, H.; and Nilsson, T. 2013 · 2013
Earlier work this paper cites.
Learning to solve arithmetic word problems with verb categorization
Hosseini, M. J.; Hajishirzi, H.; Etzioni, O.; and Kushman, N. 2014 · 2014
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W.; Yogatama, D.; Dyer, C.; and Blunsom, P. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A.; Herzig, J.; Lourie, N.; and Berant, J. 2018 · 2018
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Patel, A.; Bhattamishra, S.; and Goyal, N. 2021 · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V.; Webson, A.; Raffel, C.; Bach, S. H.; Sutawika, L.; Alyafeai, Z.; Chaffin, A.; Stiegler, A.; Scao, T. L.; Raja, A.; et al. 2021 · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z.; Wallace, E.; Feng, S.; Klein, D.; and Singh, S. 2021 · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention
Xu, Y.; Zhu, C.; Wang, S.; Sun, S.; Cheng, H.; Liu, X.; Gao, J.; He, P.; Zeng, M.; and Huang, X. 2022 · 2022
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022 · 2022
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2023 · 2023
Closest in time.
DERA: enhancing large language model completions with dialog-enabled resolving agents
Nair, V.; Schumacher, E.; Tso, G.; and Kannan, A. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Reasoning like program executors
Pi, X.; Liu, Q.; Chen, B.; Ziyadi, M.; Lin, Z.; Fu, Q.; Gao, Y.; Lou, J.-G.; and Chen, W. 2022 · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Cited alongside, same era.
Smith, S.; Patwary, M.; Norick, B.; LeGresley, P.; Rajbhandari, S.; Casper, J.; Liu, Z.; Prabhumoye, S.; Zerveas, G.; Korthikanti, V.; et al. 2022 · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2022 · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
OpenAI. 2023 · 2023
Closest in time.
Synthetic prompting: Generating chain-of-thought demonstrations for large language models
Shao, Z.; Gong, Y.; Shen, Y.; Huang, M.; Duan, N.; and Chen, W. 2023 · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N.; Cassano, F.; Labash, B.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023 · 2023
Closest in time.
Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models
Sun, J.; Luo, Y.; Gong, Y.; Lin, C.; Shen, Y.; Guo, J.; and Duan, N. 2023 · 2023
Closest in time.
Wang, Z.; Cai, S.; Liu, A.; Ma, X.; and Liang, Y. 2023 · 2023
Closest in time.
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Zhang, X.; Li, C.; Zong, Y.; Ying, Z.; He, L.; and Qiu, X. 2023 · 2023
Closest in time.
Progressive-hint prompting improves reasoning in large language models
Zheng, C.; Liu, Z.; Xie, E.; Li, Z.; and Li, Y. 2023 · 2023
Closest in time.