Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
The internal state of an llm knows when its lying
Azaria, A. and Mitchell, T · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., and Wang, H · 2023
Earlier work this paper cites.
Gemini: A Family of Highly Capable Multimodal Models, 2023
Gemini Team · 2023
Earlier work this paper cites.
Hallucinations in large multilingual translation models
Guerreiro, N. M., Alves, D. M., Waldendorf, J., Haddow, B., Birch, A., Colombo, P., and Martins, A. F · 2023
Earlier work this paper cites.
Visual programming: Compositional visual reasoning without training
Gupta, T. and Kembhavi, A · 2023
Earlier work this paper cites.
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
Hao, S., Liu, T., Wang, Z., and Hu, Z · 2023
Earlier work this paper cites.
Tool documentation enables zero-shot tool-usage with large language models
Hsieh, C.-Y., Chen, S.-A., Li, C.-L., Fujii, Y., Ratner, A., Lee, C.-Y., Krishna, R., and Pfister, T · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Earlier work this paper cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Earlier work this paper cites.
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Mündler, N., He, J., Jenko, S., and Vechev, M · 2023
Earlier work this paper cites.
Gorilla: Large Language Model Connected with Massive APIs
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Earlier work this paper cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Cited alongside, same era.
Toolformer: Language Models Can Teach Themselves to Use Tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. W.-t · 2023
Cited alongside, same era.
ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases, 2023
Tang, Q., Deng, Z., Lin, H., Han, X., Liang, Q., and Sun, L · 2023
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models
Guo, Z., Cheng, S., Wang, H., Liang, S., Qin, Y., Li, P., Liu, Z., Sun, M., and Liu, Y · 2024
Closest in time.
MetaGPT: Meta programming for a multi-agent collaborative framework
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., and Schmidhuber, J · 2024
Closest in time.
On mitigating code llm hallucinations with api documentation
Jain, N., Kwiatkowski, R., Ray, B., Ramanathan, M. K., and Kumar, V · 2024
Closest in time.
GeneGPT: augmenting large language models with domain tools for improved access to biomedical information
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Varshney, N., Yao, W., Zhang, H., Chen, J., and Yu, D · 2023
Cited alongside, same era.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2023
Cited alongside, same era.
Yang, L., Chen, H., Li, Z., Ding, X., and Wu, X · 2023
Cited alongside, same era.
Instruction tuning for large language models: A survey
Zhang, S., Dong, L., Li, X., Zhang, S., Sun, X., Wang, S., Li, J., Hu, R., Zhang, T., Wu, F., et al · 2023
Cited alongside, same era.
Verify-and-edit: A knowledge-enhanced chain-of-thought framework
Zhao, R., Li, X., Joty, S., Qin, C., and Bing, L · 2023
Cited alongside, same era.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Cited alongside, same era.
ToolQA: A Dataset for LLM Question Answering with External Tools
Zhuang, Y., Yu, Y., Wang, K., Sun, H., and Zhang, C · 2023
Cited alongside, same era.
Jin, Q., Yang, Y., Chen, Q., and Lu, Z · 2024
Closest in time.
Tool learning with foundation models, 2024
Qin, Y., Hu, S., Lin, Y., Chen, W., Ding, N., Cui, G., Zeng, Z., Huang, Y., Xiao, C., Han, C., Fung, Y. R., Su, Y., Wang, H., Qian, C., Tian, R., Zhu, K., Liang, S., Shen, X., Xu, B., Zhang, Z., Ye, Y., Li, B., Tang, Z., Yi, J., Zhu, Y., Dai, Z., Yan, L., Cong, X., Lu, Y., Zhao, W., Huang, Y., Yan, J., Han, X., Sun, X., Li, D., Phang, J., Yang, C., Wu, T., Ji, H., Liu, Z., and Sun, M · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Xu, H., Zhu, Z., Ma, D., Zhang, S., Fan, S., Chen, L., and Yu, K · 2024
Closest in time.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Closest in time.
Steptool: A step-grained reinforcement learning framework for tool learning in llms
Yu, Y., Wang, Z., Ma, W., Guo, Z., Zhan, J., Wang, S., Wu, C., Guo, Z., and Zhang, M · 2024
Closest in time.
Zhang, Y., Chen, J., Wang, J., Liu, Y., Yang, C., Shi, C., Zhu, X., Lin, Z., Wan, H., Yang, Y., et al · 2024
Closest in time.
Alignment for efficient tool calling of large language models
Xu, H., Wang, Z., Zhu, Z., Pan, L., Chen, X., Chen, L., and Yu, K · 2025
Closest in time.
Enhancing llm reliability via explicit knowledge boundary modeling
Zheng, H., Xu, H., Liu, Y., Chen, L., Fung, P., and Yu, K · 2025
Closest in time.