Fetching the paper…
Reading the bibliography…
Compound AI systems, such as agentic systems, are an emerging trend in large-scale enterprise settings, with multiple LLMs specialized for different users, tasks, and/or roles working together.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model, 2019
Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R. Radev · 2019
Earlier work this paper cites.
Understanding deep architecture with reasoning layer
Xinshi Chen, Yufei Zhang, Christoph Reisinger, and Le Song · 2020
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps, 2020
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Efficient synthesis of compact deep neural networks, 2020
Wenhan Xia, Hongxu Yin, and Niraj K. Jha · 2020
Earlier work this paper cites.
A general survey on attention mechanisms in deep learning
Gianni Brauwers and Flavius Frasincar · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
A review on the attention mechanism of deep learning
Zhaoyang Niu, Guoqiang Zhong, and Hui Yu · 2021
Earlier work this paper cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2022
Earlier work this paper cites.
Social simulacra: Creating populated prototypes for social computing systems
Joon Sung Park, Lindsay Popowski, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2022
Earlier work this paper cites.
GQA: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Earlier work this paper cites.
Distributed inference and fine-tuning of large language models over the internet, 2023
Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin Raffel · 2023
Earlier work this paper cites.
Fireact: Toward language agent fine-tuning, 2023
Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao · 2023
Earlier work this paper cites.
Punica: Multi-tenant lora serving, 2023
Lequn Chen, Zihao Ye, Yongji Wu, Danyang Zhuo, Luis Ceze, and Arvind Krishnamurthy · 2023
Earlier work this paper cites.
Mindagent: Emergent gaming interaction, 2023
Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi Vo, Zane Durante, Yusuke Noda, Zilong Zheng, Song-Chun Zhu, Demetri Terzopoulos, Li Fei-Fei, and Jianfeng Gao · 2023
Earlier work this paper cites.
Longcoder: A long-range pre-trained language model for code completion, 2023
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention, 2023
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Repobench: Benchmarking repository-level code auto-completion systems, 2023
Tianyang Liu, Canwen Xu, and Julian McAuley · 2023
Earlier work this paper cites.
Spotserve: Serving generative large language models on preemptible instances, 2023
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Earlier work this paper cites.
Artificial intelligence (ai) applications
Athanasios Valavanidis · 2023
Earlier work this paper cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Earlier work this paper cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang · 2023
Earlier work this paper cites.
The rise and potential of large language model based agents: A survey, 2023
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wensen Cheng, Qi Zhang, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, and Tao Gui · 2023
Earlier work this paper cites.
Attentionviz: A global view of transformer attention, 2023
Catherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen, Fernanda Viégas, and Martin Wattenberg · 2023
Earlier work this paper cites.
amazon/MistralLite, 2023
Yin Song and Chen Wu and Eden Duthie · 2023
Earlier work this paper cites.
Disc-lawllm: Fine-tuning large language models for intelligent legal services, 2023
Shengbin Yue, Wei Chen, Siyuan Wang, Bingxuan Li, Chenchen Shen, Shujun Liu, Yuxuan Zhou, Yao Xiao, Song Yun, Xuanjing Huang, and Zhongyu Wei · 2023
Earlier work this paper cites.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen · 2023
Earlier work this paper cites.
Llm-coordination: Evaluating and analyzing multi-agent coordination abilities in large language models, 2024
Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang · 2024
Earlier work this paper cites.
Taming throughput-latency tradeoff in llm inference with sarathi-serve, 2024
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Earlier work this paper cites.
Llama 3 model card
AI@Meta · 2024
Earlier work this paper cites.
Automated test generation to evaluate tool-augmented llms as conversational ai agents, 2024
Samuel Arcadinho, David Aparicio, and Mariana Almeida · 2024
Earlier work this paper cites.
Sustainable digitalization of business with multi-agent rag and llm
Muhammad Arslan, Saba Munawar, and Christophe Cruz · 2024
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding, 2024
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Earlier work this paper cites.
The role of ai-enhanced personalization in customer experiences
Mohammad Shafiquzzaman Bhuiyan · 2024
Earlier work this paper cites.
Coder: Issue resolving with multi-agent and task graphs, 2024
Dong Chen, Shaoxin Lin, Muhan Zeng, Daoguang Zan, Jian-Gang Wang, Anton Cheshkov, Jun Sun, Hao Yu, Guoliang Dong, Artem Aliev, Jie Wang, Xiao Cheng, Guangtai Liang, Yuchi Ma, Pan Bian, Tao Xie, and Qianxiang Wang · 2024
Earlier work this paper cites.
S-agents: Self-organizing agents in open-ended environments, 2024
Jiaqi Chen, Yuxian Jiang, Jiachen Lu, and Li Zhang · 2024
Cited alongside, same era.
When large language models meet personalization: Perspectives of challenges and opportunities
Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, et al · 2024
Cited alongside, same era.
Is bigger and deeper always better? probing llama across scales and layers, 2024
Nuo Chen, Ning Wu, Shining Liang, Ming Gong, Linjun Shou, Dongmei Zhang, and Jia Li · 2024
Cited alongside, same era.
Kvdirect: Distributed disaggregated llm inference, 2024
Shiyang Chen, Rain Jiang, Dezhi Yu, Jinlai Xu, Mengyuan Chao, Fanlong Meng, Chenyu Jiang, Wei Xu, and Hang Liu · 2024
Cited alongside, same era.
Agent-flan: Designing data and methods of effective agent tuning for large language models, 2024
Zehui Chen, Kuikun Liu, Qiuchen Wang, Wenwei Zhang, Jiangning Liu, Dahua Lin, Kai Chen, and Feng Zhao · 2024
Cited alongside, same era.
Transformers are multi-state rnns, 2024
Matanel Oren, Michael Hassid, Nir Yarden, Yossi Adi, and Roy Schwartz · 2024
Closest in time.
Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning, 2024
Rui Pan, Xiang Liu, Shizhe Diao, Renjie Pi, Jipeng Zhang, Chi Han, and Tong Zhang · 2024
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting, 2024
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Closest in time.
Predibase finetuning for customer service
Predibase · 2024
Closest in time.
Distributed Communication Package - torch.distributed
PyTorch Contributors · 2024
Closest in time.
Chatdev: Communicative agents for software development, 2024
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yihua Cheng, Kuntai Du, Jiayi Yao, and Junchen Jiang · 2024
Cited alongside, same era.
Lmcache: A kv cache sharing layer for fast distributed llm serving
LMCache Contributors · 2024
Cited alongside, same era.
A complete survey on llm-based ai chatbots, 2024
Sumit Kumar Dam, Choong Seon Hong, Yu Qiao, and Chaoning Zhang · 2024
Cited alongside, same era.
Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech, 2024
Chenpeng Du, Yiwei Guo, Hankun Wang, Yifan Yang, Zhikang Niu, Shuai Wang, Hui Zhang, Xie Chen, and Kai Yu · 2024
Cited alongside, same era.
A multi-agent conversational recommender system, 2024
Jiabao Fang, Shen Gao, Pengjie Ren, Xiuying Chen, Suzan Verberne, and Zhaochun Ren · 2024
Cited alongside, same era.
A primer on the inner workings of transformer-based language models, 2024
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà · 2024
Cited alongside, same era.
Cost-efficient large language model serving for multi-turn conversations with cachedattention, 2024
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo · 2024
Cited alongside, same era.
Closest in time.
Llama-nas: Efficient neural architecture search for large language models, 2024
Anthony Sarah, Sharath Nittur Sridhar, Maciej Szankin, and Sairam Sundaresan · 2024
Closest in time.
Small llms are weak tool learners: A multi-llm agent, 2024
Weizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan, Xiaojun Quan, Hehong Chen, Ji Zhang, and Fei Huang · 2024
Closest in time.
S-lora: Serving thousands of concurrent lora adapters, 2024
Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Fairness in serving large language models, 2024
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Agentbank: Towards generalized llm agents via fine-tuning on 50000+ interaction trajectories
Yifan Song, Weimin Xiong, Xiutian Zhao, Dawei Zhu, Wenhao Wu, Ke Wang, Cheng Li, Wei Peng, and Sujian Li · 2024
Closest in time.
Autogen 0.2 documentation - agentchat auto feedback from code execution
Microsoft AutoGen Team · 2024
Closest in time.
Fast distributed inference serving for large language models, 2024
Bingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu, Fangyue Liu, Yuanhang Sun, Gang Huang, Xuanzhe Liu, and Xin Jin · 2024
Closest in time.
dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving
Bingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun, Xuanzhe Liu, and Xin Jin · 2024
Closest in time.
Efficient multi-task llm quantization and serving for multiple lora adapters
Yifei Xia, Fangcheng Fu, Wentao Zhang, Jiawei Jiang, and Bin Cui · 2024
Closest in time.
Duoattention: Efficient long-context llm inference with retrieval and streaming heads, 2024
Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han · 2024
Closest in time.
Efficient streaming language models with attention sinks, 2024
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2024
Closest in time.
Multi-llm-agent systems: Techniques and business perspectives
Yingxuan Yang, Qiuying Peng, Jun Wang, and Weinan Zhang · 2024
Closest in time.
Cacheblend: Fast large language model serving for rag with cached knowledge fusion, 2024
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang · 2024
Closest in time.
Layer-wise importance matters: Less memory for better performance in parameter-efficient fine-tuning of large language models, 2024
Kai Yao, Penglei Gao, Lichun Li, Yuan Zhao, Xiaofeng Wang, Wei Wang, and Jianke Zhu · 2024
Closest in time.
Mammoth2: Scaling instructions from the web, 2024
Xiang Yue, Tuney Zheng, Ge Zhang, and Wenhu Chen · 2024
Closest in time.
Proagent: Building proactive cooperative agents with large language models
Ceyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang, Guanghe Li, Yihang Sun, Cheng Zhang, Zhaowei Zhang, Anji Liu, Song-Chun Zhu, Xiaojun Chang, Junge Zhang, Feng Yin, Yitao Liang, and Yaodong Yang · 2024
Closest in time.
Building cooperative embodied agents modularly with large language models, 2024
Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B. Tenenbaum, Tianmin Shu, and Chuang Gan · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving, 2024
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Manus ai agent
Monica (Butterfly Effect AI) · 2025
Closest in time.
Dynamo: A datacenter scale distributed inference serving framework
ai-dynamo · 2025
Closest in time.
Cursor: The ai code editor
Anysphere Inc · 2025
Closest in time.
Atlas: Agent tuning via learning critical steps, 2025
Zhixun Chen, Ming Li, Yuxuan Huang, Yali Du, Meng Fang, and Tianyi Zhou · 2025
Closest in time.
Github copilot
GitHub · 2025
Closest in time.
Alchemist: Towards the design of efficient online continual learning system, 2025
Yuyang Huang, Yuhan Liu, Haryadi S. Gunawi, Beibin Li, and Changho Hwang · 2025
Closest in time.
Compute or load kv cache? why not both?, 2025
Shuowei Jin, Xueshen Liu, Qingzhao Zhang, and Z. Morley Mao · 2025
Closest in time.
Azure openai service api version lifecycle
Microsoft Docs · 2025
Closest in time.
torch.cuda.stream — pytorch 2.2 documentation
PyTorch Team · 2025
Closest in time.
vLLM Production Stack: reference system for k8s-native cluster-wide deployment with community-driven performance optimization
vLLM Production Stack Team · 2025
Closest in time.
AIBrix: Cost-efficient and pluggable infrastructure components for genai inference
vLLM Project · 2025
Closest in time.