Fetching the paper…
Reading the bibliography…
Despite the impressive capabilities of large language models (LLMs), they currently exhibit two primary limitations, \textbf{\uppercase\expandafter{\romannumeral 1}}: They struggle to \textbf{autonomously solve the real world engineering problem}.
Integration of workflow and agent technology for business process management
Yan, Y., Maamar, Z., and Shen, W · 2001
Earlier work this paper cites.
Towards adaptive workflow enactment using multiagent systems
Buhler, P. A. and Vidal, J. M · 2005
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Conversational health agents: A personalized llm-powered agent framework
Abbasian, M., Azimi, I., Rahmani, A. M., and Jain, R · 2023
Earlier work this paper cites.
Metagpt: Meta programming for multi-agent collaborative framework
Hong, S., Zheng, X., Chen, J., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., et al · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
A comprehensive overview of large language models
Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
Browsergym: A reinforcement learning environment for web browsing
Research, S · 2023
Cited alongside, same era.
Devin - ai-powered collaborative teammate
AI, C · 2024
Cited alongside, same era.
22nd international conference on artificial intelligence in medicine - aime 2024
AIME · 2024
Cited alongside, same era.
Claude 3 model card october addendum
Anthropic · 2024
Cited alongside, same era.
Codeforces-contests
Codeforces · 2024
Cited alongside, same era.
Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects
Hadi, M. U., Al Tashi, Q., Shah, A., Qureshi, R., Muneer, A., Irfan, M., Zafar, A., Shaikh, M. B., Akhtar, N., Wu, J., et al · 2024
Cited alongside, same era.
Autocoder: Enhancing code large language model with \ \backslash textsc { \{ AIEV-Instruct } \}
Lei, B., Li, Y., and Chen, Q · 2024
Closest in time.
Llm critics help catch llm bugs
McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M., and Leike, J · 2024
Closest in time.
Autogen: A programming framework for agentic ai
Microsoft · 2024
Closest in time.
Chatdev: Communicative agents for software development
Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., et al · 2024
Closest in time.
Agentgpt
Team, A · 2024
Closest in time.
Aider - ai pair programming
Team, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Dang, K., et al · 2024
Cited alongside, same era.
Robin3d: Improving 3d large language model via robust instruction tuning
Kang, W., Huang, H., Shang, Y., Shah, M., and Yan, Y · 2024
Cited alongside, same era.
Macm: Utilizing a multi-agent system for condition mining in solving complex mathematical problems
Lei, B · 2024
Cited alongside, same era.
Claude 3.5 sonnet model card addendum
Anthropic, A
Cited in the paper.
Hello gpt-4o
OpenAI
Cited in the paper.
Learning to reason with llms
OpenAI
Cited in the paper.
Cursor - ai code editor
Team, C · 2024
Closest in time.
Qwen2.5-72b-instruct
Team, Q · 2024
Closest in time.
Segvg: Transferring object bounding box to segmentation for visual grounding
Kang, W., Liu, G., Shah, M., and Yan, Y · 2025
Closest in time.