Fetching the paper…
Reading the bibliography…
Existing methods fail to effectively steer Large Language Models (LLMs) between textual reasoning and code generation, leaving symbolic computing capabilities underutilized.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2022
Earlier work this paper cites.
Language models of code are few-shot commonsense learners
Madaan, A., Zhou, S., Alon, U., Yang, Y., and Neubig, G · 2022
Earlier work this paper cites.
Defining and characterizing reward gaming
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Earlier work this paper cites.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
Valmeekam, K., Olmo, A., Sreedharan, S., and Kambhampati, S · 2022
Earlier work this paper cites.
Generating sequences by learning to self-correct
Welleck, S., Lu, X., West, P., Brahman, F., Shen, T., Khashabi, D., and Choi, Y · 2022
Earlier work this paper cites.
Re3: Generating longer stories with recursive reprompting and revision
Yang, K., Tian, Y., Peng, N., and Klein, D · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Chain of code: Reasoning with a language model-augmented code emulator
Li, C., Liang, J., Zeng, A., Chen, X., Hausman, K., Sadigh, D., Levine, S., Fei-Fei, L., Xia, F., and Ichter, B · 2023
Cited alongside, same era.
Text2motion: From natural language instructions to feasible plans
Lin, K., Agia, C., Migimatsu, T., Pavone, M., and Bohg, J · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C · 2023
Meta-prompting: Enhancing language models with task-agnostic scaffolding
Suzgun, M. and Kalai, A. T · 2024
Later among the works it cites.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., and Kambhampati, S · 2024
Later among the works it cites.
Mixture-of-agents enhances large language model capabilities
Wang, J., Wang, J., Athiwaratkun, B., Zhang, C., and Zou, J · 2024
Later among the works it cites.
Learning to reason via program generation, emulation, and search
Weir, N., Khalifa, M., Qiu, L., Weller, O., and Clark, P · 2024
Later among the works it cites.
Crab: Cross-environment agent benchmark for multimodal language model agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Codeplan: Repository-level coding using llms and planning
Bairi, R., Sonwane, A., Kanade, A., Iyer, A., Parthasarathy, S., Rajamani, S., Ashok, B., and Shet, S · 2024
Cited alongside, same era.
Autotamp: Autoregressive task and motion planning with llms as translators and checkers
Chen, Y., Arkin, J., Dawson, C., Zhang, Y., Roy, N., and Fan, C · 2024
Cited alongside, same era.
PRompt optimization in multi-step tasks (PROMST): Integrating human feedback and heuristic-based sampling
Chen, Y., Arkin, J., Hao, Y., Zhang, Y., Roy, N., and Fan, C · 2024
Cited alongside, same era.
Scalable multi-robot collaboration with large language models: Centralized or decentralized systems?
Chen, Y., Arkin, J., Zhang, Y., Roy, N., and Fan, C · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Logicgame: Benchmarking rule-based reasoning abilities of large language models
Gui, J., Liu, Y., Cheng, J., Gu, X., Liu, X., Wang, H., Dong, Y., Tang, J., and Huang, M · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Xu, T., Chen, L., Wu, D.-J., Chen, Y., Zhang, Z., Yao, X., Xie, Z., Chen, Y., Liu, S., Qian, B., et al · 2024
Later among the works it cites.
Can llms reason in the wild with programs?
Yang, Y., Xiong, S., Payani, A., Shareghi, E., and Fekri, F · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, Y., Bai, H., Lin, Z., Pan, J., Tong, S., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., et al · 2024
Later among the works it cites.
Chain of preference optimization: Improving chain-of-thought reasoning in llms
Zhang, X., Du, C., Pang, T., Liu, Q., Gao, W., and Lin, M · 2024
Later among the works it cites.
Resmoe: Space-efficient compression of mixture of experts llms via residual restoration
Ai, M., Wei, T., Chen, Y., Zeng, Z., Zhao, R., Varatkar, G., Rouhani, B. D., Tang, X., Tong, H., and He, J · 2025
Closest in time.
Steering large language models between code execution and textual reasoning, 2025
Chen, Y., Jhamtani, H., Sharma, S., Fan, C., and Wang, C · 2025
Closest in time.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking
Guan, X., Zhang, L. L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F., and Yang, M · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.