Fetching the paper…
Reading the bibliography…
A diverse array of reasoning strategies has been proposed to elicit the capabilities of large language models.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Attention is turing-complete
Pérez, J., Barceló, P., and Marinkovic, J · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Fault-aware neural code rankers
Inala, J. P., Wang, C., Yang, M., Codas, A., Encarnación, M., Lahiri, S., Musuvathi, M., and Gao, J · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Generating natural language proofs with verifier-guided search
Yang, K., Deng, J., and Chen, D · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Scharli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Bousquet, O., Le, Q., and Chi, E · 2022
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Cited alongside, same era.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Cited alongside, same era.
Making language models better reasoners with step-aware verifier
Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S · 2023
Training language models with language feedback at scale
Scheurer, J., Campos, J. A., Korbak, T., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Labash, B., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Team, G · 2023
Later among the works it cites.
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deductive verification of chain-of-thought reasoning
Ling, Z., Fang, Y., Li, X., Huang, Z., Lee, M., Memisevic, R., and Su, H · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Cited alongside, same era.
The expresssive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2023
Cited alongside, same era.
Demystifying gpt self-repair for code generation
Olausson, T. X., Inala, J. P., Wang, C., Gao, J., and Solar-Lezama, A · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Improving code generation by training with natural language feedback
Chen, A., Scheurer, J., Korbak, T., Campos, J. A., Chan, J. S., Bowman, S. R., Cho, K., and Perez, E
Cited in the paper.
Introspective tips: Large language model for in-context decision making
Chen, L., Wang, L., Dong, H., Du, Y., Yan, J., Yang, F., Li, S., Zhao, P., Qin, S., Rajmohan, S., et al
Cited in the paper.
Later among the works it cites.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R., and Lim, E.-P · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Answering questions by meta-reasoning over multiple chains of thought
Yoran, O., Wolfson, T., Bogin, B., Katz, U., Deutch, D., and Berant, J · 2023
Later among the works it cites.
How language model hallucinations can snowball
Zhang, M., Press, O., Merrill, W., Liu, A., and Smith, N. A · 2023
Later among the works it cites.
Progressive-hint prompting improves reasoning in large language models
Zheng, C., Liu, Z., Xie, E., Li, Z., and Li, Y · 2023
Later among the works it cites.