Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) broaden their capabilities to manage thousands of API calls, they are confronted with complex data operations across vast datasets with significant overhead to the underlying system.
D. Marculescu, D. Stamoulis, and E. Cai, “Hardware-aware machine learning: Modeling and optimization,” in 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . IEEE, 2018, pp. 1–8
2018
Earlier work this paper cites.
D. Stamoulis, R. Ding, D. Wang, D. Lymberopoulos, B. Priyantha, J. Liu, and D. Marculescu, “Single-path nas: Designing hardware-efficient convnets in less than 4 hours,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2019, pp. 481–497
2019
Earlier work this paper cites.
C. Guo, Z. Tian, J. Tang, S. Li, Z. Wen, K. Wang, and T. Wang, “Retrieval-augmented gpt-3.5-based text-to-sql framework with sample-aware prompting and dynamic revision chain,” in International Conference on Neural Information Processing . Springer, 2023, pp. 341–356
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems , vol. 36, pp. 21 702–21 720, 2023
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Earlier work this paper cites.
H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu, “Llmlingua: Compressing prompts for accelerated inference of large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 13 358–13 376
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Zhuang, Y. Yu, K. Wang, H. Sun, and C. Zhang, “Toolqa: A dataset for llm question answering with external tools,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
Y. Qin, S. Liang et al. , “Toolllm: Facilitating large language models to master 16000+ real-world apis,” International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig, “Webarena: A realistic web environment for building autonomous agents,” in ICLR 2024 Foundation Models for Decision Making , 2024
2024
Cited alongside, same era.
M. Fore, S. Singh, and D. Stamoulis, “Geckopt: Llm system efficiency via intent-based tool selection,” in Proceedings of the Great Lakes Symposium on VLSI 2024 , 2024, pp. 353–354
2024
Closest in time.
S. Singh, M. Fore, and D. Stamoulis, “Geollm-engine: A realistic environment for building geospatial copilots,” in CVPR Workshop EarthVision , 2024
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
M. Fore, S. Singh, C. Lee, A. Pandey, A. Anastasopoulos, and D. Stamoulis, “Unlearning climate misinformation in large language models,” in Proceedings of the 1st Workshop on Natural Language Processing Meets Climate Change (ClimateNLP 2024) . ACL, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried, “Visualwebarena: Evaluating multimodal agents on realistic visual web tasks,” ACL , 2024
2024
Cited alongside, same era.
S. Singh, M. Fore, and D. Stamoulis, “Evaluating tool-augmented agents in remote sensing platforms,” in ICLR 2024 2nd Workshop on ML for Remote Sensing , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
LangChain docs, “Langchain llm caching,” 2024, accessed: May 2024
2024
Closest in time.