Fetching the paper…
Reading the bibliography…
LLM Ensemble -- which involves the comprehensive use of multiple large language models (LLMs), each aimed at handling user queries during downstream inference, to benefit from their individual strengths -- has gained substantial attention recently.
Sentence-bert: Sentence embeddings using siamese bert-networks
N Reimers · 2019
Earlier work this paper cites.
Structured probabilistic end-to-end learning from crowds
Zhijun Chen, Huimin Wang, Hailong Sun, Pengpeng Chen, Tao Han, Xudong Liu, and Jie Yang · 2021
Earlier work this paper cites.
Wrench: A comprehensive benchmark for weak supervision
Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang, and Alexander Ratner · 2021
Earlier work this paper cites.
Adversarial learning from crowds
Pengpeng Chen, Hailong Sun, Yongqiang Yang, and Zhijun Chen · 2022
Earlier work this paper cites.
Relative representations enable zero-shot latent space communication
Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, and Emanuele Rodolà · 2022
Earlier work this paper cites.
Model cascading: Towards jointly improving efficiency and accuracy of nlp systems
Neeraj Varshney and Chitta Baral · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Automix: Automatically mixing language models
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, et al · 2023
Earlier work this paper cites.
Frugalgpt: How to use large language models while reducing cost and improving performance
Lingjiao Chen, Matei Zaharia, and James Zou · 2023
Earlier work this paper cites.
Black-box data poisoning attacks on crowdsourcing
Pengpeng Chen, Yongqiang Yang, Dingqi Yang, Hailong Sun, Zhijun Chen, and Peng Lin · 2023
Earlier work this paper cites.
Neural-hidden-crf: A robust weakly-supervised sequence labeler
Zhijun Chen, Hailong Sun, Wanhao Zhang, Chunyi Xu, Qianren Mao, and Pengpeng Chen · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Earlier work this paper cites.
Ensemble-instruct: Generating instruction-tuning data with a heterogeneous mixture of lms
Young-Suk Lee, Md Arafat Sultan, Yousef El-Kurdi, Tahira Naseem Asim Munawar, Radu Florian, Salim Roukos, and Ramón Fernandez Astudillo · 2023
Earlier work this paper cites.
Cache & distil: Optimising api calls to large language models
Guillem Ramírez, Matthias Lindemann, Alexandra Birch, and Ivan Titov · 2023
Earlier work this paper cites.
Large language model routing with benchmark datasets
Tal Shnitzer, Anthony Ou, Mirian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson, and Mikhail Yurochkin · 2023
Earlier work this paper cites.
Getting more out of mixture of language model reasoning experts
Chenglei Si, Weijia Shi, Chen Zhao, Luke Zettlemoyer, and Jordan Boyd-Graber · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Ecoassistant: Using llm assistant more affordably and accurately
Jieyu Zhang, Ranjay Krishna, Ahmed H Awadallah, and Chi Wang · 2023
Earlier work this paper cites.
A unified approach to routing and cascading for llms
Jasper Dekoninck, Maximilian Baader, and Martin Vechev · 2024
Earlier work this paper cites.
Hybrid llm: Cost-efficient and quality-aware query routing
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks VS Lakshmanan, and Ahmed Hassan Awadallah · 2024
Cited alongside, same era.
Bayesian calibration of win rate estimation with llm evaluators
Yicheng Gao, Gonghan Xu, Zhe Wang, and Arman Cohan · 2024
Cited alongside, same era.
Smoothie: Label free language model routing
Neel Guha, Mayee F Chen, Trevor Chow, Ishan S Khare, and Christopher Re · 2024
Cited alongside, same era.
Satya Kesav Gundabathula and Sriram R Kolar · 2024
Cited alongside, same era.
Language model cascades: Token-level uncertainty and beyond
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, and Sanjiv Kumar · 2024
Routoo: Learning to route to large language models effectively
Alireza Mohammadshahi, Arshad Rafiq Shaikh, and Majid Yazdani · 2024
Later among the works it cites.
Adaptive selection for homogeneous tools: An instantiation in the rag scenario
Feiteng Mu, Yong Jiang, Liwen Zhang, Chu Liu, Wenjie Li, Pengjun Xie, and Fei Huang · 2024
Later among the works it cites.
Metallm: A high-performant and cost-efficient dynamic framework for wrapping llms
Quang H Nguyen, Duy C Hoang, Juliette Decugis, Saurav Manchanda, Nitesh V Chawla, and Khoa D Doan · 2024
Later among the works it cites.
Routellm: Learning to route llms with preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamic ensemble reasoning for llm experts
Jinwu Hu, Yufeng Wang, Shuhai Zhang, Kai Zhou, Guohao Chen, Yu Hu, Bin Xiao, and Mingkui Tan · 2024
Cited alongside, same era.
Routerbench: A benchmark for multi-llm routing system
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay · 2024
Cited alongside, same era.
Ensemble learning for heterogeneous large language models with deep parallel collaboration
Yichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang, Hui Wang, Ting Liu, and Bing Qin · 2024
Cited alongside, same era.
Collaborative decoding of critical tokens for boosting factuality of large language models
Lifeng Jin, Baolin Peng, Linfeng Song, Haitao Mi, Ye Tian, and Dong Yu · 2024
Cited alongside, same era.
When does confidence-based cascade deferral suffice?
Wittawat Jitkrittum, Neha Gupta, Aditya K Menon, Harikrishna Narasimhan, Ankit Rawat, and Sanjiv Kumar · 2024
Cited alongside, same era.
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye · 2024
Cited alongside, same era.
Purifying large language models by ensembling a small language model
Tianlin Li, Qian Liu, Tianyu Pang, Chao Du, Qing Guo, Yang Liu, and Min Lin · 2024
Cited alongside, same era.
Sungjin Park, Xiao Liu, Yeyun Gong, and Edward Choi · 2024
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Marija Šakota, Maxime Peyrard, and Robert West · 2024
Later among the works it cites.
Pickllm: Context-aware rl-assisted large language model routing
Dimitrios Sikeridis, Dennis Ramdass, and Pranay Pareek · 2024
Later among the works it cites.
Harnessing the power of multiple minds: Lessons learned from llm routing
KV Srivatsa, Kaushal Kumar Maurya, and Ekaterina Kochmar · 2024
Later among the works it cites.
Tensoropera router: A multi-model router for efficient llm inference
Dimitris Stripelis, Zijian Hu, Jipeng Zhang, Zhaozhuo Xu, Alay Dilipbhai Shah, Han Jin, Yuhang Yao, Salman Avestimehr, and Chaoyang He · 2024
Later among the works it cites.
Llm-topla: Efficient llm ensemble by maximising diversity
Selim Tekin, Fatih Ilhan, Tiansheng Huang, Sihao Hu, and Ling Liu · 2024
Later among the works it cites.
Bench-coe: a framework for collaboration of experts from benchmark
Yuanshuai Wang, Xingjian Zhang, Jinkun Zhao, Siwei Wen, Peilin Feng, Shuhao Liao, Lei Huang, and Wenjun Wu · 2024
Later among the works it cites.
Bridging the gap between different vocabularies for llm ensemble
Yangyifan Xu, Jinliang Lu, and Jiajun Zhang · 2024
Later among the works it cites.
Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao · 2024
Later among the works it cites.
Determine-then-ensemble: Necessity of top-k union for large language model ensembling
Yuxuan Yao, Han Wu, Mingyang Liu, Sichun Luo, Xiongwei Han, Jie Liu, Zhijiang Guo, and Linqi Song · 2024
Later among the works it cites.
Yao-Ching Yu, Chun-Chih Kuo, Ziqi Ye, Yu-Cheng Chang, and Yueh-Se Li · 2024
Later among the works it cites.
Large language model cascades with mixture of thought representations for cost-efficient reasoning
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao · 2024
Later among the works it cites.
Eagle: Efficient training-free router for multi-llm inference
Zesen Zhao, Shuowei Jin, and Z Morley Mao · 2024
Later among the works it cites.
Llm bandit: Cost-efficient llm generation via preference-conditioned dynamic routing
Yang Li · 2025
Closest in time.
Hit the sweet spot! span-level ensemble for large language models
Yangyifan Xu, Jianghao Chen, Junhong Wu, and Jiajun Zhang · 2025
Closest in time.
Citer: Collaborative inference for efficient large language model decoding with token-level routing
Wenhao Zheng, Yixiao Chen, Weitong Zhang, Souvik Kundu, Yun Li, Zhengzhong Liu, Eric P Xing, Hongyi Wang, and Huaxiu Yao · 2025
Closest in time.