Fetching the paper…
Reading the bibliography…
Self-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs.
Meta-learning without memorization
M. Yin, G. Tucker, M. Zhou, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Earlier work this paper cites.
Offline rl for natural language generation with implicit language q learning
C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Earlier work this paper cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
A. F. Akyürek, E. Akyürek, A. Madaan, A. Kalyan, P. Clark, D. Wijaya, and N. Tandon · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
X. Chen, M. Lin, N. Schärli, and D. Zhou · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou · 2023
Earlier work this paper cites.
Language models can solve computer tasks
G. Kim, P. Baldi, and S. McAleer · 2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2023
Cited alongside, same era.
Is self-repair a silver bullet for code generation?
T. X. Olausson, J. P. Inala, C. Wang, J. Gao, and A. Solar-Lezama · 2023
Cited alongside, same era.
L. Pan, M. Saxon, W. Xu, D. Nathani, X. Wang, and W. Y. Wang · 2023
Cited alongside, same era.
Refiner: Reasoning feedback on intermediate representations
D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and B. Faltings · 2023
Cited alongside, same era.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Next: Teaching large language models to reason about code execution
A. Ni, M. Allamanis, A. Cohan, Y. Deng, K. Shi, C. Sutton, and P. Yin · 2024
Closest in time.
Recursive introspection: Teaching language model agents how to self-improve
Y. Qu, T. Zhang, N. Garg, and A. Kumar · 2024
Closest in time.
Multi-turn reinforcement learning from preference human feedback
L. Shani, A. Rosenberg, A. Cassel, O. Lang, D. Calandriello, A. Zipori, H. Noga, O. Keller, B. Piot, I. Szpektor, et al · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Closest in time.
Beyond human data: Scaling self-training for problem-solving with language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Shinn, B. Labash, and A. Gopinath · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, et al · 2023
Cited alongside, same era.
Generating sequences by learning to self-correct
S. Welleck, X. Lu, P. West, F. Brahman, T. Shen, D. Khashabi, and Y. Choi · 2023
Cited alongside, same era.
Selfee: Iterative self-revising llm empowered by self-feedback generation
S. Ye, Y. Jo, D. Kim, S. Kim, H. Hwang, and M. Seo · 2023
Cited alongside, same era.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, A. Üstün, and S. Hooker · 2024
Cited alongside, same era.
Stop regressing: Training value functions via classification for scalable deep rl
J. Farebrother, J. Orbay, Q. Vuong, A. A. Taïga, Y. Chebotar, T. Xiao, A. Irpan, S. Levine, P. S. Castro, A. Faust, et al · 2024
Cited alongside, same era.
Reference-free monolithic preference optimization with odds ratio
J. Hong, N. Lee, and J. Thorne · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Cited alongside, same era.
A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, A. Kumar, A. Alemi, A. Rizkowsky, A. Nova, B. Adlam, B. Bohnet, G. Elsayed, H. Sedghi, I. Mordatch, I. Simpson, I. Gur, J. Snoek, J. Pennington, J. Hron, K. Kenealy, K. Swersky, K. Mahajan, L. Culp, L. Xiao, M. L. Bileschi, N. Constant, R. Novak, R. Liu, T. Warkentin, Y. Qian, Y. Bansal, E. Dyer, B. Neyshabur, J. Sohl-Dickstein, and N. Fiedel · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Closest in time.
Codegemma: Open code models based on gemma
C. Team · 2024
Closest in time.
Llms cannot find reasoning errors, but can correct them given the error location
G. Tyen, H. Mansoor, V. Cărbune, Y. P. Chen, and T. Mak · 2024
Closest in time.
Building math agents with multi-turn iterative preference learning
W. Xiong, C. Shi, J. Shen, A. Rosenberg, Z. Qin, D. Calandriello, M. Khalman, R. Joshi, B. Piot, M. Saleh, et al · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?
S. Yang, E. Gribovskaya, N. Kassner, M. Geva, and S. Riedel · 2024
Closest in time.
Physics of language models: Part 2.2, how to learn from mistakes on grade-school math problems, 2024
T. Ye, Z. Xu, Y. Li, and Z. Allen-Zhu · 2024
Closest in time.
Small language models need strong verifiers to self-correct reasoning
Y. Zhang, M. Khalifa, L. Logeswaran, J. Kim, M. Lee, H. Lee, and L. Wang · 2024
Closest in time.
Natural plan: Benchmarking llms on natural language planning
H. S. Zheng, S. Mishra, H. Zhang, X. Chen, M. Chen, A. Nova, L. Hou, H.-T. Cheng, Q. V. Le, E. H. Chi, et al · 2024
Closest in time.
Archer: Training language model agents via hierarchical multi-turn rl
Y. Zhou, A. Zanette, J. Pan, S. Levine, and A. Kumar · 2024
Closest in time.