Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are transforming artificial intelligence, evolving into task-oriented systems capable of autonomous planning and execution.
Conversational health agents: A personalized llm-powered agent framework, 2023
M. Abbasian, I. Azimi, A. M. Rahmani, and R. Jain · 2023
Earlier work this paper cites.
Benchmarking llm powered chatbots: Methods and metrics, 2023
D. Banerjee, P. Singh, A. Avadhanam, and S. Srivastava · 2023
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, TimBrooks, J. Wang, L. Li, LongOuyang, JuntangZhuang, JoyceLee, YufeiGuo, WesamManassra, PrafullaDhariwal, CaseyChu, YunxinJiao, and A. Ramesh · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
N. Ding, Y. Chen, B. Xu, Y. Qin, S. Hu, Z. Liu, M. Sun, and B. Zhou · 2023
Earlier work this paper cites.
Botchat: Evaluating llms’ capabilities of having multi-turn dialogues, 2023
H. Duan, J. Wei, C. Wang, H. Liu, Y. Fang, S. Zhang, D. Lin, and K. Chen · 2023
Earlier work this paper cites.
Ragas: Automated evaluation of retrieval-augmented generation
S. Es, J. James, L. Espinosa-Anke, and S. Schockaert · 2023
Earlier work this paper cites.
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li · 2023
Earlier work this paper cites.
Unnatural instructions: Tuning language models with (almost) no human labor
O. Honovich, T. Scialom, O. Levy, and T. Schick · 2023
Earlier work this paper cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. Canton-Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2023
Earlier work this paper cites.
Explore-instruct: Enhancing domain-specific instruction coverage through active exploration
F. Wan, X. Huang, T. Yang, X. Quan, W. Bi, and S. Shi · 2023
Earlier work this paper cites.
Let’s synthesize step by step: Iterative dataset synthesis with large language models by extrapolating errors from small models
R. Wang, W. Zhou, and M. Sachan · 2023
Earlier work this paper cites.
Wizardlm: Empowering large language models to follow complex instructions, 2023
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Cited alongside, same era.
R. Xu, H. Cui, Y. Yu, X. Kan, W. Shi, Y. Zhuang, W. Jin, J. C. Ho, and C. J. Yang · 2023
Cited alongside, same era.
Decoding data quality via synthetic corruptions: Embedding-guided pruning of code data
Y. Yang, A. K. Singh, M. Elhoushi, A. Mahmoud, K. Tirumala, F. Gloeckle, B. Rozière, C. Wu, A. S. Morcos, and N. Ardalani · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Z. Yuan, H. Yuan, C. Li, G. Dong, C. Tan, and C. Zhou · 2023
Cited alongside, same era.
Automated test generation to evaluate tool-augmented llms as conversational ai agents
On llms-driven synthetic data generation, curation, and evaluation: A survey, 2024
L. Long, R. Wang, R. Xiao, J. Zhao, X. Ding, G. Chen, and H. Wang · 2024
Later among the works it cites.
Evaluation and continual improvement for an enterprise ai assistant, 2024
A. V. Maharaj, K. Qian, U. Bhattacharya, S. Fang, H. Galatanu, M. Garg, R. Hanessian, N. Kapoor, K. Russell, S. Vaithyanathan, and Y. Li · 2024
Later among the works it cites.
Evaluating large language models as agents in the clinic
N. Mehandru, B. Y. Miao, E. R. Almaraz, M. Sushil, A. J. Butte, and A. Alaa · 2024
Later among the works it cites.
“ask me anything”: How comcast uses llms to assist agents in real time
S. Rome, T. Chen, R. Tang, L. Zhou, and F. Ture · 2024
Later among the works it cites.
Improving text embeddings with large language models
L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Arcadinho, D. Aparício, and M. S. C. Almeida · 2024
Cited alongside, same era.
Beyond prompts: Dynamic conversational benchmarking of large language models
D. Castillo-Bolado, J. Davidson, F. Gray, and M. Rosa · 2024
Cited alongside, same era.
Large language model agent in financial trading: A survey, 2024
H. Ding, Y. Li, J. Wang, and H. Chen · 2024
Cited alongside, same era.
From local to global: A graph rag approach to query-focused summarization
D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, and J. Larson · 2024
Cited alongside, same era.
Crmarena: Understanding the capacity of llm agents to perform professional crm tasks in realistic environments, 2024
K.-H. Huang, A. Prabhakar, S. Dhawan, Y. Mao, H. Wang, S. Savarese, C. Xiong, P. Laban, and C.-S. Wu · 2024
Cited alongside, same era.
MT-eval: A multi-turn capabilities evaluation benchmark for large language models
W.-C. Kwan, X. Zeng, Y. Jiang, Y. Wang, L. Li, L. Shang, X. Jiang, Q. Liu, and K.-F. Wong · 2024
Cited alongside, same era.
Intent-based prompt calibration: Enhancing prompt optimization with synthetic boundary cases, 2024
E. Levi, E. Brosh, and M. Friedmann · 2024
Cited alongside, same era.
Tradingagents: Multi-agents llm financial trading framework, 2024
Y. Xiao, E. Sun, D. Luo, and W. Wang · 2024
Later among the works it cites.
Eduagent: Generative student agents in learning, 2024
S. Xu, X. Zhang, and L. Qin · 2024
Later among the works it cites.
Finrobot: An open-source ai agent platform for financial applications using large language models, 2024
H. Yang, B. Zhang, N. Wang, C. Guo, X. Zhang, L. Lin, J. Wang, T. Zhou, M. Guan, R. Zhang, and C. D. Wang · 2024
Later among the works it cites.
Content Knowledge Identification with Multi-agent Large Language Models (LLMs)
K. Yang, Y. Chu, T. Darwin, A. Han, H. Li, H. Wen, Y. Copur-Gencturk, J. Tang, and H. Liu · 2024
Later among the works it cites.
τ-bench: A benchmark for tool-agent-user interaction in real-world domains
S. Yao, N. Shinn, P. Razavi, and K. Narasimhan · 2024
Later among the works it cites.
Curate: Benchmarking personalised alignment of conversational ai assistants
L. Alberts, B. Ellis, A. Lupu, and J. Foerster · 2025
Closest in time.