Fetching the paper…
Reading the bibliography…
We address the problem of code generation from multi-turn execution feedback.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Archer: Training language model agents via hierarchical multi-turn rl
Zhou, Y., Zanette, A., Pan, J., Levine, S., and Kumar, A · 2002
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search, 2017
Anthony, T., Tian, Z., and Barber, D · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, S · 2021
Earlier work this paper cites.
Codet: Code generation with generated tests, 2022
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., and Chen, W · 2022
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning, 2022
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., and Hoi, S. C. H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Cited alongside, same era.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Mapcoder: Multi-agent code generation for competitive problem solving, 2024
Islam, M. A., Ali, M. E., and Parvez, M. R · 2024
Later among the works it cites.
NExt: Teaching large language models to reason about code execution
Ni, A., Allamanis, M., Cohan, A., Deng, Y., Shi, K., Sutton, C., and Yin, P · 2024
Later among the works it cites.
Recursive introspection: Teaching language model agents how to self-improve, 2024
Qu, Y., Zhang, T., Garg, N., and Kumar, A · 2024
Later among the works it cites.
Code generation with alphacodium: From prompt engineering to flow engineering, 2024
Ridnik, T., Kredo, D., and Friedman, I · 2024
Later among the works it cites.
Rewarding progress: Scaling automated process verifiers for llm reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S · 2023
Cited alongside, same era.
Execution-based code generation using deep reinforcement learning, 2023
Shojaee, P., Jain, A., Tipirneni, S., and Reddy, C. K · 2023
Cited alongside, same era.
Generating sequences by learning to self-correct
Welleck, S., Lu, X., West, P., Brahman, F., Shen, T., Khashabi, D., and Choi, Y · 2023
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Grounding large language models in interactive environments with online reinforcement learning, 2024
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y · 2024
Cited alongside, same era.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2024
Cited alongside, same era.
Better than your teacher: Llm agents that learn from privileged ai feedback, 2024
Choudhury, S. and Sodhi, P · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Setlur, A., Nagpal, C., Fisch, A., Geng, X., Eisenstein, J., Agarwal, R., Agarwal, A., Berant, J., and Kumar, A · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning, 2024
Zhai, Y., Bai, H., Lin, Z., Pan, J., Tong, S., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., and Levine, S · 2024
Later among the works it cites.
Rest-mcts*: Llm self-training via process reward guided tree search, 2024
Zhang, D., Zhoubian, S., Hu, Z., Yue, Y., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Commit0: Library generation from scratch
Zhao, W., Jiang, N., Lee, C., Chiu, J. T., Cardie, C., Gallé, M., and Rush, A. M · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs, 2024
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., and Sheng, Y · 2024
Later among the works it cites.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking, 2025
Guan, X., Zhang, L. L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F., and Yang, M · 2025
Closest in time.