Fetching the paper…
Reading the bibliography…
The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback).
Codearena: Inspecting and improving code quality metrics using minecraft
S. Baars and S. Meester · 2019
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, et al · 2021
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Santacoder: don’t reach for the stars!
L. B. Allal, R. Li, D. Kocetkov, C. Mou, C. Akiki, C. M. Ferrandis, N. Muennighoff, M. Mishra, A. Gu, M. Dey, et al · 2023
Earlier work this paper cites.
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, et al · 2023
Earlier work this paper cites.
Wizardcoder: Empowering code large language models with evol-instruct
Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, et al · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
Critique-out-loud reward models
Z. Ankner, M. Paul, B. Cui, J. D. Chang, and P. Ammanabrolu · 2024
Cited alongside, same era.
Introducing claude
Anthropic · 2024
Cited alongside, same era.
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
G. Bai, J. Liu, X. Bu, Y. He, J. Liu, Z. Zhou, Z. Lin, W. Su, T. Ge, B. Zheng, et al · 2024
Cited alongside, same era.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Llm critics help catch llm bugs
N. McAleese, R. M. Pokorny, J. F. C. Uribe, E. Nitishinskaya, M. Trebacz, and J. Leike · 2024
Later among the works it cites.
Hello gpt-4o
OpenAI · 2024
Later among the works it cites.
Hellobench: Evaluating long text generation capabilities of large language models
H. Que, F. Duan, L. He, Y. Mou, W. Zhou, J. Liu, W. Rong, Z. M. Wang, J. Yang, G. Zhang, et al · 2024
Later among the works it cites.
A critical evaluation of ai feedback for aligning large language models
A. Sharma, S. Keh, E. Mitchell, C. Finn, K. Arora, and T. Kollar · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
S. Tan, S. Zhuang, K. Montgomery, W. Y. Tang, A. Cuadron, C. Wang, R. A. Popa, and I. Stoica · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li, et al · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, et al · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. I. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Cited alongside, same era.
Critiquellm: Towards an informative critique generation model for evaluation of large language model generation
P. Ke, B. Wen, A. Feng, X. Liu, X. Lei, J. Cheng, S. Wang, A. Zeng, Y. Dong, H. Wang, et al · 2024
Cited alongside, same era.
Criticbench: Benchmarking llms for critique-correct reasoning, 2024
Z. Lin, Z. Gou, T. Liang, R. Luo, H. Liu, and Y. Yang · 2024
Cited alongside, same era.
E2-LLM: Efficient and extreme length extension of large language models
J. Liu, Z. ZhiqiBai, Y. Zhang, C. Zhang, Y. YuangZh, G. Zhang, J. JiakaiWang, H. Que, Y. Chen, W. Su, T. Ge, J. Fu, W. Chen, and B. Zheng · 2024
Cited alongside, same era.
Training language models to critique with multi-agent feedback
T. Lan, W. Zhang, C. Lyu, S. Li, C. Xu, H. Huang, D. Lin, X.-L. Mao, and K. Chen
Cited in the paper.
Later among the works it cites.
The llama 3 herd of models
L. Team · 2024
Later among the works it cites.
Supercorrect: Supervising and correcting language models with error-driven insights
L. Yang, Z. Yu, T. Zhang, M. Xu, J. E. Gonzalez, B. Cui, and S. Yan · 2024
Later among the works it cites.
Mammoth2: Scaling instructions from the web
X. Yue, T. Zheng, G. Zhang, and W. Chen · 2024
Later among the works it cites.
Generative verifiers: Reward modeling as next-token prediction
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2024
Later among the works it cites.
Realcritic: Towards effectiveness-driven evaluation of language model critiques
Z. Tang, Z. Li, Z. Xiao, T. Ding, R. Sun, B. Wang, D. Liu, F. Huang, T. Liu, B. Yu, et al · 2025
Closest in time.
MTU-bench: A multi-granularity tool-use benchmark for large language models
P. Wang, Y. Wu, N. Wang, J. Liu, X. Song, Z. Peng, K. Deng, C. Zhang, JiakaiWang, J. Peng, G. Zhang, H. Guo, Z. Zhang, W. Su, and B. Zheng · 2025
Closest in time.