Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,” Biometrika , vol. 39, no. 3-4, pp. 324–345, 1952
1952
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Shi, A. Karpathy et al. , “World of bits: An open-domain platform for web-based agents,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp. 3135–3144. [Online]. Available: https://proceedings.mlr.press/v70/shi17a.html
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Kočiskỳ, J. Schwarz et al. , “The narrativeqa reading comprehension challenge,” Transactions of the Association for Computational Linguistics , vol. 6, pp. 317–328, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M.-A. Côté, A. Kádár et al. , “Textworld: A learning environment for text-based games,” in Computer Games: 7th Workshop, CGW 2018, Held in Conjunction with the 27th International Conference on Artificial Intelligence, IJCAI 2018, Stockholm, Sweden, July 13, 2018, Revised Selected Papers 7 . Springer, 2019, pp. 41–75
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hafner, “Benchmarking the spectrum of agent capabilities,” arXiv preprint arXiv:2109.06780 , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Geva, D. Khashabi et al. , “Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 346–361, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
C. Rawles, A. Li et al. , “Androidinthewild: A large-scale dataset for android device control,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 708–59 728, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Lee, K. Lee et al. , “Qasa: advanced question answering on scientific articles,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 036–19 052
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Valmeekam, M. Marquez et al. , “Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change,” Advances in Neural Information Processing Systems , vol. 36, pp. 38 975–38 987, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Li et al. , “Api-bank: A comprehensive benchmark for tool-augmented llms,” EMNLP , 2023. [Online]. Available: https://aclanthology.org/2023.emnlp-main.187/
2023
Earlier work this paper cites.
R. Team, “Restgpt: Connecting llms with real-world restful apis,” RestGPT , 2023. [Online]. Available: https://restgpt.github.io/
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Alahi et al. , “Smarteval: Evaluation system for descriptive answers in examinations using iot-enabled technologies and artificial intelligence,” Sensors , vol. 23, no. 11, p. 5206, 2023. [Online]. Available: https://doi.org/10.3390/s23115206
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. A. Team, “Galileo ai agent leaderboard,” 2023. [Online]. Available: https://huggingface.co/spaces/galileo-ai/agent-leaderboard
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Carta, C. Romac et al. , “Grounding large language models in interactive environments with online reinforcement learning,” in International Conference on Machine Learning . PMLR, 2023, pp. 3676–3713
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
G. Team, “Berkeley function calling leaderboard v3 (aka berkeley tool calling leaderboard v3),” Gorilla , 2024. [Online]. Available: https://gorilla.cs.berkeley.edu/leaderboard.html
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
N. Team, “Nexusraven-13b, a new sota open-source llm for function calling,” Nexusflow , 2024. [Online]. Available: https://github.com/nexusflowai/NexusRaven
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Yu, B. Shen et al. , “Codereval: A benchmark of pragmatic code generation with generative pre-trained models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
“Swe arena: An open evaluation platform for automated software engineering,” 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
K. Basu et al. , “A comprehensive corpora for training and benchmarking api llms,” ACL , 2024. [Online]. Available: https://aclanthology.org/2024.acl-long.694/
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Guo et al. , “Towards stable large-scale benchmarking on tool learning of large language models,” ACL Findings , 2024. [Online]. Available: https://aclanthology.org/2024.findings-acl.664/
2024
Later among the works it cites.
K. Basu et al. , “A benchmark for evaluating llms on nested sequences of api calls,” OpenReview , 2024. [Online]. Available: https://openreview.net/forum?id=r7staQknbI
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Yu et al. , “Mindagent: Emergent gaming interaction,” in Findings of the Association for Computational Linguistics: NAACL 2024 , 2024, pp. 200–210. [Online]. Available: https://aclanthology.org/2024.findings-naacl.200/
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C.-K. Wu, Z. R. Tam et al. , “Streambench: Towards benchmarking continuous improvement of language agents,” Advances in Neural Information Processing Systems , vol. 37, pp. 107 039–107 063, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Renze and E. Guven, “The benefits of a concise chain of thought on problem-solving in large language models,” in 2024 2nd International Conference on Foundation and Large Language Models (FLLM) . IEEE, Nov. 2024, p. 476–483. [Online]. Available: http://dx.doi.org/10.1109/FLLM63129.2024.10852493
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W.-L. Chiang, L. Zheng et al. , “Chatbot arena: An open platform for evaluating llms by human preference,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Kara, F. Faisal, and S. Nath, “Waber: Evaluating reliability and efficiency of web agents with existing benchmarks,” in ICLR 2025 Workshop on Foundation Models in the Wild
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J.-t. Huang, E. J. Li et al. , “Competing large language models in multi-agent gaming environments,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
H. Kokel, M. Katz et al. , “Acpbench: Reasoning about action, change, and planning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 25, 2025, pp. 26 559–26 568
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
T. Team, “Tooleyes: Fine-grained evaluation for tool learning capabilities of large language models,” COLING , 2025. [Online]. Available: https://aclanthology.org/2025.coling-main.12/
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Stroebl, S. Kapoor, and A. Narayanan, “Hal: A holistic agent leaderboard for centralized and reproducible agent evaluation,” https://github.com/princeton-pli/hal-harness , 2025
2025
Closest in time.
N. Yekollu, A. Bohra, A. Chirumamilla et al. , “Agent arena: A platform for evaluating and comparing llm agents,” 2025. [Online]. Available: https://www.agent-arena.com/
2025
Closest in time.
2025
Closest in time.
Anthropic, “Claude ai,” 2025, accessed: 2025-04-27. [Online]. Available: https://claude.ai/
2025
Closest in time.
2025
Closest in time.
Asterisk, “Deepclaude: Combining deepseek r1’s reasoning with claude’s creativity and code generation,” 2025, accessed: 2025-04-27. [Online]. Available: https://deepclaude.com/
2025
Closest in time.