Fetching the paper…
Reading the bibliography…
The availability of a wide range of large language models (LLMs) embedded in various agentic systems has significantly increased the potential of model selection strategies to improve the cost-performance tradeoff.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Cascadebert: Accelerating inference of pre-trained language models via calibrated complete models cascade
Li, L., Lin, Y., Chen, D., Ren, S., Li, P., Zhou, J., and Sun, X · 2021
Earlier work this paper cites.
Efficient online ML API selection for multi-label classification tasks
Chen, L., Zaharia, M., and Zou, J · 2022
Earlier work this paper cites.
Babybear: Cheap inference triage for expensive language models
Khalili, L., You, Y., and Bohannon, J · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D · 2022
Earlier work this paper cites.
Model cascading: Towards jointly improving efficiency and accuracy of NLP systems
Varshney, N. and Baral, C · 2022
Earlier work this paper cites.
Frugalgpt: How to use large language models while reducing cost and improving performance
Chen, L., Zaharia, M., and Zou, J · 2023
Earlier work this paper cites.
Tryage: Real-time, intelligent routing of user prompts to large language models
Hari, S. N. and Thomson, M · 2023
Earlier work this paper cites.
Exploring the benefits of training expert language models over instruction tuning
Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
When does confidence-based cascade deferral suffice?
Jitkrittum, W., Gupta, N., Menon, A. K., Narasimhan, H., Rawat, A. S., and Kumar, S · 2023
Cited alongside, same era.
When does confidence-based cascade deferral suffice?
Jitkrittum, W., Gupta, N., Menon, A. K., Narasimhan, H., Rawat, A. S., and Kumar, S · 2023
Cited alongside, same era.
Automix: Automatically mixing language models
Madaan, A., Aggarwal, P., Anand, A., Potharaju, S. P., Mishra, S., Zhou, P., Gupta, A., Rajagopal, D., Kappaganthu, K., Yang, Y., Upadhyay, S., Mausam, and Faruqui, M · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Large language model routing with benchmark datasets
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M · 2023
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Closest in time.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Closest in time.
Optllm: Optimal assignment of queries to large language models
Liu, Y., Zhang, H., Miao, Y., Le, V., and Li, Z · 2024
Closest in time.
Routing to the expert: Efficient reward-guided ensemble of large language models
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J · 2024
Closest in time.
Metallm: A high-performant and cost-efficient dynamic framework for wrapping llms
Nguyen, Q. H., Hoang, D. C., Decugis, J., Manchanda, S., Chawla, N. V., and Doan, K. D · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamic voting for efficient reasoning in large language models
Xue, M., Liu, D., Lei, W., Ren, X., Yang, B., Xie, J., Zhang, Y., Peng, D., and Lv, J · 2023
Cited alongside, same era.
Data shunt: Collaboration of small and large models for lower costs and better performance
Chen, D., Zhuang, Y., Zhang, S., Liu, J., Dong, S., and Tang, S · 2024
Cited alongside, same era.
Learning to route with confidence tokens
Chuang, Y., Zhou, H., Sarma, P. K., Gopalan, P., Boccio, J., Bolouki, S., and Hu, X · 2024
Cited alongside, same era.
Learning how hard to think: Input-adaptive allocation of LM computation
Damani, M., Shenfeld, I., Peng, A., Bobu, A., and Andreas, J · 2024
Cited alongside, same era.
Hybrid LLM: cost-efficient and quality-aware query routing
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Rühle, V., Lakshmanan, L. V. S., and Awadallah, A. H · 2024
Cited alongside, same era.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Cited alongside, same era.
Language model cascades: Token-level uncertainty and beyond
Gupta, N., Narasimhan, H., Jitkrittum, W., Rawat, A. S., Menon, A. K., and Kumar, S · 2024
Cited alongside, same era.
Closest in time.
Mixeval: Deriving wisdom of the crowd from LLM benchmark mixtures
Ni, J., Xue, F., Yue, X., Deng, Y., Shah, M., Jain, K., Neubig, G., and You, Y · 2024
Closest in time.
Online cascade learning for efficient inference over streams
Nie, L., Ding, Z., Hu, E., Jermaine, C. M., and Chaudhuri, S · 2024
Closest in time.
Routellm: Learning to route llms with preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Closest in time.
Domain-Aware LLM Routing During Generation
Pichlmeier, J., Ross, P., and Luckow, A · 2024
Closest in time.
Optimising calls to large language models with uncertainty-based two-tier selection
Ramírez, G., Birch, A., and Titov, I · 2024
Closest in time.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Sakota, M., Peyrard, M., and West, R · 2024
Closest in time.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., and Chen, W · 2024
Closest in time.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement
Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., Lu, K., Xue, M., Lin, R., Liu, T., Ren, X., and Zhang, Z · 2024
Closest in time.