Fetching the paper…
Reading the bibliography…
With the rapid growth in the number of Large Language Models (LLMs), there has been a recent interest in LLM routing, or directing queries to the cheapest LLM that can deliver a suitable response.
Stacked generalization
D. H. Wolpert · 1992
Earlier work this paper cites.
Experiments with a new boosting algorithm
Y. Freund, R. E. Schapire, et al · 1996
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
J. H. Friedman · 2001
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
J.-Y. Audibert and A. B. Tsybakov · 2007
Earlier work this paper cites.
MS MARCO: A human-generated MAchine reading COmprehension dataset, 2017
T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng · 2017
Earlier work this paper cites.
Marginal Singularity, and the Benefits of Labels in Covariate-Shift
S. Kpotufe and G. Martinet · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Transfer Learning for Nonparametric Classification: Minimax Rate and Adaptive Classifier
T. T. Cai and H. Wei · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
The TechQA dataset
V. Castelli, R. Chakravarti, S. Dana, A. Ferritto, R. Florian, M. Franz, D. Garg, D. Khandelwal, S. McCarley, M. McCawley, M. Nasr, L. Pan, C. Pendus, J. Pitrelli, S. Pujar, S. Roukos, A. Sakrajda, A. Sil, R. Uceda-Sosa, T. Ward, and R. Zhang · 2020
Earlier work this paper cites.
FrugalML: How to use ML prediction APIs more accurately and cheaply
L. Chen, M. Zaharia, and J. Zou · 2020
Earlier work this paper cites.
COVID-QA: A question answering dataset for COVID-19
T. Möller, A. Reina, R. Jayakumar, and M. Pietsch · 2020
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang · 2021
Earlier work this paper cites.
Question answering over electronic devices: A new benchmark dataset and a multi-task learning based QA framework
A. Nandy, S. Sharma, S. Maddhashiya, K. Sachdeva, P. Goyal, and N. Ganguly · 2021
Earlier work this paper cites.
TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance
F. Zhu, W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua · 2021
Earlier work this paper cites.
Efficient online ml api selection for multi-label classification tasks, 2022
L. Chen, M. Zaharia, and J. Zou · 2022
Cited alongside, same era.
Minimax optimal approaches to the label shift problem in non-parametric settings
S. Maity, Y. Sun, and M. Banerjee · 2022
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang, X. Ren, and B. Y. Lin · 2023
Cited alongside, same era.
Hagrid: A human-llm collaborative dataset for generative information-seeking with attribution, 2023
E. Kamalloo, A. Jafari, X. Zhang, N. Thakur, and J. Lin · 2023
Cited alongside, same era.
Automix: Automatically mixing language models
A. Madaan, P. Aggarwal, A. Anand, S. P. Potharaju, S. Mishra, P. Zhou, A. Gupta, D. Rajagopal, K. Kappaganthu, Y. Yang, et al · 2023
Expertqa: Expert-curated questions and attributed answers, 2024
C. Malaviya, S. Lee, S. Chen, E. Sieber, M. Yatskar, and D. Roth · 2024
Later among the works it cites.
Large language models: A survey
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Later among the works it cites.
A. Myrzakhan, S. M. Bsharat, and Z. Shen · 2024
Later among the works it cites.
Mixeval: Deriving wisdom of the crowd from llm benchmark mixtures
J. Ni, F. Xue, X. Yue, Y. Deng, M. Shah, K. Jain, G. Neubig, and Y. You · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2023
Cited alongside, same era.
Delucionqa: Detecting hallucinations in domain-specific question answering
M. Sadat, Z. Zhou, L. Lange, J. Araki, A. Gundroo, B. Wang, R. Menon, M. Parvez, and Z. Feng · 2023
Cited alongside, same era.
Large language model routing with benchmark datasets
T. Shnitzer, A. Ou, M. Silva, K. Soule, Y. Sun, J. Solomon, N. Thompson, and M. Yurochkin · 2023
Cited alongside, same era.
Musr: Testing the limits of chain-of-thought with multistep soft reasoning
Z. Sprague, X. Ye, K. Bostrom, S. Chaudhuri, and G. Durrett · 2023
Cited alongside, same era.
Fusing models with complementary expertise
H. Wang, F. M. Polo, Y. Sun, S. Kundu, E. Xing, and M. Yurochkin · 2023
Cited alongside, same era.
Gpt-4 technical report, 2024
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2024
Cited alongside, same era.
FrugalGPT: How to use large language models while reducing cost and improving performance
L. Chen, M. Zaharia, and J. Zou · 2024
Cited alongside, same era.
I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica · 2024
Later among the works it cites.
Musr: Testing the limits of chain-of-thought with multistep soft reasoning, 2024
Z. Sprague, X. Ye, K. Bostrom, S. Chaudhuri, and G. Durrett · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024
Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen · 2024
Later among the works it cites.
M. Yue, J. Zhao, M. Zhang, L. Du, and Z. Yao · 2024
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
M. Šakota, M. Peyrard, and R. West · 2024
Later among the works it cites.
Ragbench: Explainable benchmark for retrieval-augmented generation systems, 2025
R. Friel, M. Belyi, and A. Sanyal · 2025
Closest in time.
Rorf: Routing on random forests, 2023
D. Jain, T.-Y. Tung, and T. H. Kofman · 2025
Closest in time.
Metallm: A high-performant and cost-efficient dynamic framework for wrapping llms, 2025
Q. H. Nguyen, T. Dao, D. C. Hoang, J. Decugis, S. Manchanda, N. V. Chawla, and K. D. Doan · 2025
Closest in time.
Openai text-embedding-3-small model, 2023
OpenAI · 2025
Closest in time.
Openhermes 2.5, 2023
Teknium · 2025
Closest in time.