Fetching the paper…
Reading the bibliography…
Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide accurate judgments and actionable suggestions.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Markov chains and stochastic stability
Meyn, S. P. and Tweedie, R. L · 2012
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Earlier work this paper cites.
Learning by distilling context
Snell, C., Klein, D., and Zhong, R · 2022
Earlier work this paper cites.
Generating sequences by learning to self-correct
Welleck, S., Lu, X., West, P., Brahman, F., Shen, T., Khashabi, D., and Choi, Y · 2022
Earlier work this paper cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Akyürek, A. F., Akyürek, E., Madaan, A., Kalyan, A., Clark, P., Wijaya, D., and Tandon, N · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2023
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback, 2023
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2023
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Earlier work this paper cites.
Critic: Large language models can self-correct with tool-interactive critiquing
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2023
Earlier work this paper cites.
Taco: Topics in algorithmic code generation dataset
Li, R., Fu, J., Zhang, B.-W., Huang, T., Sun, Z., Lyu, C., Liu, G., Jin, Z., and Li, G · 2023
Cited alongside, same era.
Debate helps supervise unreliable experts
Michael, J., Mahdi, S., Rein, D., Petty, J., Dirani, J., Padmakumar, V., and Bowman, S. R · 2023
Cited alongside, same era.
Pan, L., Saxon, M., Xu, W., Nathani, D., Wang, X., and Wang, W. Y · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
Salmon: Self-alignment with principle-following reward models
Sun, Z., Shen, Y., Zhang, H., Zhou, Q., Chen, Z., Cox, D. D., Yang, Y., and Gan, C · 2023
Next: Teaching large language models to reason about code execution
Ni, A., Allamanis, M., Cohan, A., Deng, Y., Shi, K., Sutton, C., and Yin, P · 2024
Later among the works it cites.
Spontaneous reward hacking in iterative self-refinement
Pan, J., He, H., Bowman, S. R., and Feng, S · 2024
Later among the works it cites.
Archon: An architecture search framework for inference-time techniques
Saad-Falcon, J., Lafuente, A. G., Natarajan, S., Maru, N., Todorov, H., Guha, E., Buchanan, E. K., Chen, M., Guha, N., Ré, C., et al · 2024
Later among the works it cites.
Bond: Aligning llms with best-of-n distillation
Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Vieillard, N., Ramé, A., Shariari, B., Perrin, S., Friesen, A., Cideron, G., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Shepherd: A critic for language model generation
Wang, T., Yu, P., Tan, X. E., O’Brien, S., Pasunuru, R., Dwivedi-Yu, J., Golovneva, O., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A · 2023
Cited alongside, same era.
Retroformer: Retrospective large language agents with policy gradient optimization
Yao, W., Heinecke, S., Niebles, J. C., Liu, Z., Feng, Y., Xue, L., Murthy, R., Chen, Z., Zhang, J., Arpit, D., et al · 2023
Cited alongside, same era.
Critique-out-loud reward models
Ankner, Z., Paul, M., Cui, B., Chang, J. D., and Ammanabrolu, P · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Natural language reinforcement learning
Feng, X., Wan, Z., Fu, H., Liu, B., Yang, M., Koushik, G. A., Hu, Z., Wen, Y., and Wang, J · 2024
Cited alongside, same era.
Deliberative alignment: Reasoning enables safer language models
Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Heylar, A., Dias, R., Vallone, A., Ren, H., Wei, J., et al · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Lu, K., et al · 2024
Cited alongside, same era.
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Later among the works it cites.
Hybridflow: A flexible and efficient rlhf framework
Sheng, G., Zhang, C., Ye, Z., Wu, X., Zhang, W., Zhang, R., Peng, Y., Lin, H., and Wu, C · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
A survey of neural code intelligence: Paradigms, advances and beyond
Sun, Q., Chen, Z., Xu, F., Cheng, K., Ma, C., Yin, Z., Wang, J., Han, C., Zhu, R., Yuan, S., et al · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Tan, S., Zhuang, S., Montgomery, K., Tang, W. Y., Cuadron, A., Wang, C., Popa, R. A., and Stoica, I · 2024
Later among the works it cites.
Enhancing llm reasoning via critique models with test-time and training-time supervision
Xi, Z., Yang, D., Huang, J., Tang, J., Li, G., Ding, Y., He, W., Hong, B., Do, S., Zhan, W., et al · 2024
Later among the works it cites.
Llava-critic: Learning to evaluate multimodal models
Xiong, T., Wang, X., Guo, D., Ye, Q., Fan, H., Gu, Q., Huang, H., and Li, C · 2024
Later among the works it cites.
Improving reward models with synthetic critiques
Ye, Z., Greenlee-Scott, F., Bartolo, M., Blunsom, P., Campos, J. A., and Gallé, M · 2024
Later among the works it cites.
Self-generated critiques boost reward modeling for language models
Yu, Y., Chen, Z., Zhang, A., Tan, L., Zhu, C., Pang, R. Y., Qian, Y., Wang, X., Gururangan, S., Zhang, C., et al · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
What makes large language models reason in (multi-turn) code generation?
Zheng, K., Decugis, J., Gehring, J., Cohen, T., Negrevergne, B., and Synnaeve, G · 2024
Later among the works it cites.
Ldb: A large language model debugger via verifying runtime execution step-by-step
Zhong, L., Wang, Z., and Shang, J · 2024
Later among the works it cites.