Fetching the paper…
Reading the bibliography…
Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A.; and Terry, M. E. 1952 · 1952
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning
LaValle, S. 1998 · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V.; and Tsitsiklis, J. 1999 · 1999
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L.; and Szepesvári, C. 2006 · 2006
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C. B.; Powley, E.; Whitehouse, D.; Lucas, S. M.; Cowling, P. I.; Rohlfshagen, P.; Tavener, S.; Perez, D.; Samothrakis, S.; and Colton, S. 2012 · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S.; Chen, H.; Yang, J.; and Narasimhan, K. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Fireact: Toward language agent fine-tuning
Chen, B.; Shu, C.; Shareghi, E.; Collier, N.; Narasimhan, K.; and Yao, S. 2023 · 2023
Cited alongside, same era.
Pangu-agent: A fine-tunable generalist agent with structured reasoning
Christianos, F.; Papoudakis, G.; Zimmer, M.; Coste, T.; Wu, Z.; Chen, J.; Khandelwal, K.; Doran, J.; Feng, X.; Liu, J.; et al. 2023 · 2023
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X.; Wan, Z.; Wen, M.; Wen, Y.; Zhang, W.; and Wang, J. 2023 · 2023
Cited alongside, same era.
Reasoning with Language Model is Planning with World Model
Hao, S.; Gu, Y.; Ma, H.; Hong, J.; Wang, Z.; Wang, D.; and Hu, Z. 2023 · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2023 · 2023
Tree Search for Language Model Agents
Koh, J. Y.; McAleer, S.; Fried, D.; and Salakhutdinov, R. 2024 · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Chandu, K.; Dziri, N.; Kumar, S.; Zick, T.; Choi, Y.; et al. 2024 · 2024
Closest in time.
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Luo, L.; Liu, Y.; Liu, R.; Phatale, S.; Lara, H.; Li, Y.; Shu, L.; Zhu, Y.; Meng, L.; Sun, J.; et al. 2024 · 2024
Closest in time.
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Ma, C.; Zhang, J.; Zhu, Z.; Yang, C.; Yang, Y.; Jin, Y.; Lan, Z.; Kong, L.; and He, J. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023 · 2023
Cited alongside, same era.
Decomposition enhances reasoning via self-evaluation guided decoding
Xie, Y.; Kawaguchi, K.; Zhao, Y.; Zhao, X.; Kan, M.-Y.; He, J.; and Xie, Q. 2023 · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z.; Yuan, H.; Li, C.; Dong, G.; Tan, C.; and Zhou, C. 2023 · 2023
Cited alongside, same era.
Agenttuning: Enabling generalized agent abilities for llms
Zeng, A.; Liu, M.; Lu, R.; Wang, B.; Liu, X.; Dong, Y.; and Tang, J. 2023 · 2023
Cited alongside, same era.
AlphaMath Almost Zero: process Supervision without process
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024a
Cited in the paper.
Pang, J.-C.; Wang, P.; Li, K.; Chen, X.-H.; Xu, J.; Zhang, Z.; and Yu, Y. 2024 · 2024
Closest in time.
From r to Q*: Your Language Model is Secretly a Q-Function
Rafailov, R.; Hejna, J.; Park, R.; and Finn, C. 2024 · 2024
Closest in time.
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Song, Y.; Yin, D.; Yue, X.; Huang, J.; Li, S.; and Lin, B. Y. 2024 · 2024
Closest in time.
A survey on large language model based autonomous agents
Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. 2024 · 2024
Closest in time.
Asymptotics of language model alignment
Yang, J. Q.; Salamatian, S.; Sun, Z.; Suresh, A. T.; and Beirami, A. 2024 · 2024
Closest in time.
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Zhai, Y.; Bai, H.; Lin, Z.; Pan, J.; Tong, S.; Zhou, Y.; Suhr, A.; Xie, S.; LeCun, Y.; Ma, Y.; et al. 2024 · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H.; Feng, G.; Xiong, W.; Zhao, L.; He, D.; Bian, J.; and Wang, L. 2024 · 2024
Closest in time.