Fetching the paper…
Reading the bibliography…
Large language models, employed as multiple agents that interact and collaborate with each other, have excelled at solving complex tasks.
Neural architecture search with bayesian optimisation and optimal transport
K. Kandasamy, W. Neiswanger, J. Schneider, B. Poczos, and E. P. Xing · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
H. Liu, K. Simonyan, and Y. Yang · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
X. Ho, A.-K. Duong Nguyen, S. Sugawara, and A. Aizawa · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Interpretable neural architecture search via bayesian optimisation with weisfeiler-lehman kernels
B. Ru, X. Wan, X. Dong, and M. Osborne · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
MuSiQue: Multihop questions via single-hop question composition
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal · 2022
Earlier work this paper cites.
On redundancy and diversity in cell-based neural architecture search
X. Wan, B. Ru, P. M. Esperança, and Z. Li · 2022
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2023
Earlier work this paper cites.
Automatic prompt optimization with “gradient descent” and beam search
R. Pryzant, D. Iter, J. Li, Y. Lee, C. Zhu, and M. Zeng · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Earlier work this paper cites.
Neural architecture search: Insights from 1000 papers
C. White, M. Safari, R. Sukthanker, B. Ru, T. Elsken, A. Zela, D. Dey, and F. Hutter · 2023
Earlier work this paper cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2023
Cited alongside, same era.
Survival of the most influential prompts: Efficient black-box prompt search via clustering and pruning
H. Zhou, X. Wan, I. Vulić, and A. Korhonen · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
LongBench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y. Dong, J. Tang, and J. Li · 2024
Cited alongside, same era.
ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs
J. Chen, S. Saha, and M. Bansal · 2024
Cited alongside, same era.
Universal self-consistency for large language models
AutoAct: Automatic agent learning from scratch for QA via self-planning
S. Qiao, N. Zhang, R. Fang, Y. Luo, W. Zhou, Y. Jiang, C. Lv, and H. Chen · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. P. Lillicrap, J. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, I. Antonoglou, R. Anil, S. Borgeaud, A. M. Dai, K. Millican, E. Dyer, M. Glaese, T. Sottiaux, B. Lee, F. Viola, M. Reynolds, Y. Xu, J. Molloy, J. Chen, M. Isard, P. Barham, T. Hennigan, R. McIlroy, M. Johnson, J. Schalkwyk, E. Collins, E. Rutherford, E. Moreira, K. Ayoub, M. Goel, C. Meyer, G. Thornton, Z. Yang, H. Michalewski, Z. Abbas, N. Schucher, A. Anand, R. Ives, J. Keeling, K. Lenc, S. Haykal, S. Shakeri, P. Shyam, A. Chowdhery, R. Ring, S. Spencer, E. Sezener, and et al · 2024
Later among the works it cites.
Archon: An architecture search framework for inference-time techniques
J. Saad-Falcon, A. G. Lafuente, S. Natarajan, N. Maru, H. Todorov, E. Guha, E. K. Buchanan, M. Chen, N. Guha, C. Ré, et al · 2024
Later among the works it cites.
Agentsquare: Automatic llm agent search in modular design space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen, R. Aksitov, U. Alon, J. Ren, K. Xiao, P. Yin, S. Prakash, C. Sutton, X. Wang, and D. Zhou · 2024
Cited alongside, same era.
Trace is the next autodiff: Generative optimization with rich feedback, execution traces, and LLMs
C.-A. Cheng, A. Nie, and A. Swaminathan · 2024
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2024
Cited alongside, same era.
Ds-agent: Automated data science by empowering large language models with case-based reasoning, 2024
S. Guo, C. Deng, Y. Wen, H. Chen, Y. Chang, and J. Wang · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Cited alongside, same era.
Debating with more persuasive LLMs leads to more truthful answers
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez · 2024
Cited alongside, same era.
DSPy: Compiling declarative language model calls into state-of-the-art pipelines
O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. V. A, S. Haq, A. Sharma, T. T. Joshi, H. Moazam, H. Miller, M. Zaharia, and C. Potts · 2024
Cited alongside, same era.
Y. Shang, Y. Li, K. Zhao, L. Ma, J. Liu, F. Xu, and Y. Li · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2024
Later among the works it cites.
On the brittle foundations of react prompting for agentic large language models
M. Verma, S. Bhambri, and S. Kambhampati · 2024
Later among the works it cites.
Teach better or show smarter? on instructions and exemplars in automatic prompt optimization
X. Wan, R. Sun, H. Nakhost, and S. O. Arik · 2024
Later among the works it cites.
Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key?
Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song · 2024
Later among the works it cites.
Agentless: Demystifying llm-based software engineering agents
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang · 2024
Later among the works it cites.
Large language models as optimizers
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen · 2024
Later among the works it cites.
Exploring collaboration mechanisms for LLM agents: A social psychology view
J. Zhang, X. Xu, N. Zhang, R. Liu, B. Hooi, and S. Deng · 2024
Later among the works it cites.
Agent-pro: Learning to evolve via policy-level reflection and optimization
W. Zhang, K. Tang, H. Wu, M. Wang, Y. Shen, G. Hou, Z. Tan, P. Li, Y. Zhuang, and W. Lu · 2024
Later among the works it cites.
Fairer preferences elicit improved human-aligned large language model judgments
H. Zhou, X. Wan, Y. Liu, N. Collier, I. Vulić, and A. Korhonen · 2024
Later among the works it cites.
GPTSwarm: Language agents as optimizable graphs
M. Zhuge, W. Wang, L. Kirsch, F. Faccio, D. Khizbullin, and J. Schmidhuber · 2024
Later among the works it cites.
Embodied agent interface: Benchmarking llms for embodied decision making, 2025
M. Li, S. Zhao, Q. Wang, K. Wang, Y. Zhou, S. Srivastava, C. Gokmen, T. Lee, L. E. Li, R. Zhang, W. Liu, P. Liang, L. Fei-Fei, J. Mao, and J. Wu · 2025
Closest in time.
Agentic retrieval-augmented generation: A survey on agentic rag
A. Singh, A. Ehtesham, S. Kumar, and T. T. Khoei · 2025
Closest in time.
Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments
H. Su, R. Sun, J. Yoon, P. Yin, T. Yu, and S. Ö. Arık · 2025
Closest in time.
Optimizing generative ai by backpropagating language model feedback
M. Yuksekgonul, F. Bianchi, J. Boen, S. Liu, P. Lu, Z. Huang, C. Guestrin, and J. Zou · 2025
Closest in time.
Multi-agent architecture search via agentic supernet
G. Zhang, L. Niu, J. Fang, K. Wang, L. Bai, and X. Wang · 2025
Closest in time.