Fetching the paper…
Reading the bibliography…
LLM routers aim to balance quality and cost of generation by classifying queries and routing them to a cheaper or more expensive LLM depending on their complexity.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,” Biometrika , vol. 39, no. 3/4, 1952
1952
Earlier work this paper cites.
F. Jelinek, “Interpolated estimation of Markov source parameters from sparse data,” 1980. [Online]. Available: https://api.semanticscholar.org/CorpusID:61012010
1980
Earlier work this paper cites.
N. Dalvi, P. Domingos, Mausam, S. Sanghai, and D. Verma, “Adversarial classification,” in Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining , 2004
2004
Earlier work this paper cites.
D. Lowd and C. Meek, “Adversarial learning,” in ACM International Conference on Knowledge Discovery in Data Mining (SIGKDD) , 2005
2005
Earlier work this paper cites.
2013
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in International Conference on Learning Representations (ICLR) , 2016
2016
Earlier work this paper cites.
N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami, “The limitations of deep learning in adversarial settings,” in IEEE European symposium on security and privacy (EuroS&P) , 2016
2016
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
F. Tramèr, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart, “Stealing machine learning models via prediction APIs,” in USENIX Security Symposium , 2016
2016
Earlier work this paper cites.
N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami, “Practical black-box attacks against machine learning,” in Proceedings of the 2017 ACM on Asia conference on computer and communications security , 2017
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf , 2019
2019
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,” Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Earlier work this paper cites.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning (ICML) , 2022
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research (JMLR) , 2022
2022
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in NeurIPS ML Safety Workshop , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection,” in ACM AISec , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“OpenAI and others seek new path to smarter AI as current methods hit limitations,” https://www.reuters.com/technology/artificial-intelligence/openai-rivals-seek-new-path-smarter-ai-current-methods-hit-limitations-2024-11-11 , published: 2024-11-15
2024
Later among the works it cites.
“OpenAI, Google and Anthropic are struggling to build more advanced AI,” https://www.bloomberg.com/news/articles/2024-11-13/openai-google-and-anthropic-are-struggling-to-build-more-advanced-ai?sref=CrGXSfHu , published: 2024-11-13
2024
Later among the works it cites.
“OpenAI shifts strategy as rate of ‘GPT’ AI improvements slows,” https://www.theinformation.com/articles/openai-shifts-strategy-as-rate-of-gpt-ai-improvements-slows , published: 2024-11-9
2024
Later among the works it cites.
“What is a control plane?” https://www.ibm.com/think/topics/control-plane , published: 2024-10-31
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Jiang, X. Ren, and B. Y. Lin, “LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Narayanan Hari and M. Thomson, “Tryage: Real-time, intelligent routing of user prompts to large language models,” arXiv e-prints , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Schulhoff, J. Pinto, A. Khan, L.-F. Bouchard, C. Si, S. Anati, V. Tagliabue, A. Kost, C. Carnahan, and J. Boyd-Graber, “Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition,” in EMNLP , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Teknium, “Openhermes 2.5: An open dataset of synthetic data for generalist LLM assistants,” 2023. [Online]. Available: https://huggingface.co/datasets/teknium/OpenHermes-2.5
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, and I. Stoica, “Chatbot arena: An open platform for evaluating LLMs by human preference,” in Forty-first International Conference on Machine Learning (ICML) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. Rühle, L. V. Lakshmanan, and A. H. Awadallah, “Hybrid LLM: Cost-efficient and quality-aware query routing,” in International Conference on Learning Representations (ICLR) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C.-H. Lee, H. Cheng, and M. Ostendorf, “OrchestraLLM: Efficient orchestration of language models for dialogue state tracking,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Šakota, M. Peyrard, and R. West, “Fly-swat or cannon? cost-effective language model choice via meta-modeling,” in Proceedings of the 17th ACM International Conference on Web Search and Data Mining , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Yue, J. Zhao, M. Zhang, L. Du, and Z. Yao, “Large language model cascades with mixture of thought representations for cost-efficient reasoning,” in International Conference on Learning Representations (ICLR) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zhang, N. Carlini, and D. Ippolito, “Effective prompt extraction from language models,” in First Conference on Language Modeling , 2024
2024
Later among the works it cites.