Fetching the paper…
Reading the bibliography…
Large Language Model (LLM)-based systems, i.e.
Iulia Turc, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 1908
Earlier work this paper cites.
“DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf · 1910
Earlier work this paper cites.
“Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.”
Ralph Bradley and Milton. Terry · 1952
Earlier work this paper cites.
“Learning Automata - A Survey”
Kumpati. Narendra and M… Thathachar · 1974
Earlier work this paper cites.
“Ratings of Chess Players Past and Present”
Arpad. Elo · 1978
Earlier work this paper cites.
“Adaptive Mixtures of Local Experts”
Robert. Jacobs, Michael. Jordan, Steven. Nowlan and Geoffrey. Hinton · 1991
Earlier work this paper cites.
“Reinforcement learning: a survey”
Leslie Kaelbling, Michael. Littman and Andrew. Moore · 1996
Earlier work this paper cites.
“ROUGE: A Package for Automatic Evaluation of Summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
“Preference Learning”
Johannes Fürnkranz and Eyke Hüllermeier · 2012
Earlier work this paper cites.
“Temporal-Difference Learning”
Richard. Sutton and Andrew. Barto · 2014
Earlier work this paper cites.
“Proximal Policy Optimization Algorithms”, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov · 2017
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Contrastive Representation Learning: A Framework and Review”
Phuc Le-Khac, Graham Healy and Alan Smeaton · 2020
Earlier work this paper cites.
“BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov and Luke Zettlemoyer · 2020
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive NLP tasks”
Patrick Lewis et al · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter. Liu · 2020
Earlier work this paper cites.
“Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics”
Prajjwal Bhargava, Aleksandr Drozd and Anna Rogers · 2021
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt · 2021
Earlier work this paper cites.
“BARTScore: Evaluating Generated Text as Text Generation”
Weizhe Yuan, Graham Neubig and Pengfei Liu · 2021
Earlier work this paper cites.
“A Robustly Optimized BERT Pre-training Approach with Post-training”
Liu Zhuang, Lin Wayne, Shi Ya and Zhao Jun · 2021
Earlier work this paper cites.
“A Literature Survey of Recent Advances in Chatbots”
Guendalina Caldarini, Sardar Jaf and Kenneth McGarry · 2022
Earlier work this paper cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2022
Earlier work this paper cites.
“Aligning artificial intelligence with climate change mitigation”
Lynn. Kaack, Priya. Donti, Emma Strubell, George Kamiya, Felix Creutzig and David Rolnick · 2022
Earlier work this paper cites.
“Emergent Abilities of Large Language Models”
Jason Wei et al · 2022
Earlier work this paper cites.
Yuefeng Zhang · 2022
Earlier work this paper cites.
Jinze Bai et al · 2023
Earlier work this paper cites.
“FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”, 2023
Lingjiao Chen, Matei Zaharia and James Zou · 2023
Earlier work this paper cites.
“A Closer Look into Using Large Language Models for Automatic Evaluation”
Cheng-Han Chiang and Hung-yi Lee · 2023
Earlier work this paper cites.
“Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning”
Radosvet Desislavov, Fernando Martínez-Plumed and José Hernández-Orallo · 2023
Earlier work this paper cites.
“Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models”, 2023
Surya Hari and Matt Thomson · 2023
Earlier work this paper cites.
Pengcheng He, Jianfeng Gao and Weizhu Chen · 2023
Earlier work this paper cites.
“DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing”
Pengcheng He, Jianfeng Gao and Weizhu Chen · 2023
Earlier work this paper cites.
“Exploring the benefits of training expert language models over instruction tuning”
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee and Minjoon Seo · 2023
Earlier work this paper cites.
“LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion”
Dongfu Jiang, Xiang Ren and Bill Lin · 2023
Earlier work this paper cites.
“Evaluating Open-Domain Question Answering in the Era of Large Language Models”
Ehsan Kamalloo, Nouha Dziri, Charles Clarke and Davood Rafiei · 2023
Earlier work this paper cites.
“A Review of AI-Driven Conversational Chatbots Implementation Methodologies and Challenges (1999–2022)”
Chien-Chang Lin, Anna.. Huang and Stephen.. Yang · 2023
Earlier work this paper cites.
“HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu and Yueting Zhuang · 2023
Earlier work this paper cites.
“Getting MoRE out of Mixture of Language Model Reasoning Experts”
Chenglei Si, Weijia Shi, Chen Zhao, Luke Zettlemoyer and Jordan Boyd-Graber · 2023
Cited alongside, same era.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”, 2023
Hugo Touvron et al · 2023
Cited alongside, same era.
“EcoAssistant: Using LLM Assistant More Affordably and Accurately”, 2023
Jieyu Zhang, Ranjay Krishna, Ahmed. Awadallah and Chi Wang · 2023
Cited alongside, same era.
“AutoMix: Automatically Mixing Language Models”
Pranjal Aggarwal et al · 2024
Cited alongside, same era.
“Derby LLM : Évaluation comparative des approches RAG et fine-tuning”
Christophe Bouvard, Mathieu Ciancone, Antoine Gourru and Marion Schaeffer · 2024
Cited alongside, same era.
“Performance Characterization of Expert Router for Scalable LLM Inference”, 2024
Josef Pichlmeier, Philipp Ross and Andre Luckow · 2024
Later among the works it cites.
“Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection”
Guillem Ramírez, Alexandra Birch and Ivan Titov · 2024
Later among the works it cites.
“Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling”
Marija Sakota, Maxime Peyrard and Robert West · 2024
Later among the works it cites.
“Large Language Model Routing with Benchmark Datasets”, 2024
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson and Mikhail Yurochkin · 2024
Later among the works it cites.
“PickLLM: Context-Aware RL-Assisted Large Language Model Routing”, 2024
Dimitrios Sikeridis, Dennis Ramdass and Pranay Pareek · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing”
Ziwei Chai et al · 2024
Cited alongside, same era.
“A Survey on Evaluation of Large Language Models”
Yupeng Chang et al · 2024
Cited alongside, same era.
“RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models”
Shuhao Chen, Weisen Jiang, Baijiong Lin, James Kwok and Yu Zhang · 2024
Cited alongside, same era.
“Learning to Route with Confidence Tokens”, 2024
Yu-Neng Chuang, Helen Zhou, Prathusha Sarma, Parikshit Gopalan, John Boccio, Sara Bolouki and Xia Hu · 2024
Cited alongside, same era.
“Scaling Instruction-Finetuned Language Models”
Hyung Chung et al · 2024
Cited alongside, same era.
“A Complete Survey on LLM-based AI Chatbots”, 2024
Sumit Dam, Choong Hong, Yu Qiao and Chaoning Zhang · 2024
Cited alongside, same era.
“Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing”
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks.. Lakshmanan and Ahmed Awadallah · 2024
Cited alongside, same era.
Later among the works it cites.
“MoDEM: Mixture of Domain Expert Models”
Toby Simonds, Kemal Kurniawan and Jey Lau · 2024
Later among the works it cites.
“Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing”
Kv Srivatsa, Kaushal Maurya and Ekaterina Kochmar · 2024
Later among the works it cites.
“TensorOpera Router: A Multi-Model Router for Efficient LLM Inference”
Dimitris Stripelis, Zhaozhuo Xu, Zijian Hu, Alay Shah, Han Jin, Yuhang Yao, Jipeng Zhang, Tong Zhang, Salman Avestimehr and Chaoyang He · 2024
Later among the works it cites.
“A survey on effective invocation methods of massive LLM services”, 2024
Can Wang, Bolin Zhang, Dianbo Sui, Zhiying Tu, Xiaoyu Liu and Jiabao Kang · 2024
Later among the works it cites.
“Fusing Models with Complementary Expertise”
Hongyi Wang, Felipe Polo, Yuekai Sun, Souvik Kundu, Eric Xing and Mikhail Yurochkin · 2024
Later among the works it cites.
“Improving Text Embeddings with Large Language Models”
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder and Furu Wei · 2024
Later among the works it cites.
“Bench-CoE: a Framework for Collaboration of Experts from Benchmark”, 2024
Yuanshuai Wang, Xingjian Zhang, Jinkun Zhao, Siwei Wen, Peilin Feng, Shuhao Liao, Lei Huang and Wenjun Wu · 2024
Later among the works it cites.
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin and Mei Han · 2024
Later among the works it cites.
“Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He and Bryan Hooi · 2024
Later among the works it cites.
“Qwen2 Technical Report”, 2024
An Yang et al · 2024
Later among the works it cites.
“FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets”
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim and Minjoon Seo · 2024
Later among the works it cites.
“Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning”
Murong Yue, Jie Zhao, Min Zhang, Liang Du and Ziyu Yao · 2024
Later among the works it cites.
Liang Zhang, Katherine Jijo, Spurthi Setty, Eden Chung, Fatima Javid, Natan Vidra and Tommy Clifford · 2024
Later among the works it cites.
“Eagle: Efficient Training-Free Router for Multi-LLM Inference”, 2024
Zesen Zhao, Shuowei Jin and Z. Mao · 2024
Later among the works it cites.
“LoraRetriever: Input-Aware LoRA Retrieval and Composition for Mixed Tasks in the Wild”
Ziyu Zhao, Leilei Gan, Guoyin Wang, Wangchunshu Zhou, Hongxia Yang, Kun Kuang and Fei Wu · 2024
Later among the works it cites.
Adarsh Behera, Jaya Champati, Roberto Morabito, Sasu Tarkoma and James Gross · 2025
Closest in time.
“A Survey on Mixture of Experts in Large Language Models”
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim and Jiayi Huang · 2025
Closest in time.
“A Survey on Collaborative Mechanisms Between Large and Small Language Models”, 2025
Yi Chen, JiaHao Zhao and HaoHao Han · 2025
Closest in time.
“Harnessing Multiple Large Language Models: A Survey on LLM Ensemble”, 2025
Zhijun Chen, Jingzheng Li, Pengpeng Chen, Zhuoran Li, Kai Sun, Yuankai Luo, Qianren Mao, Dingqi Yang, Hailong Sun and Philip. Yu · 2025
Closest in time.
“A Unified Approach to Routing and Cascading for LLMs”, 2025
Jasper Dekoninck, Maximilian Baader and Martin Vechev · 2025
Closest in time.
“GraphRouter: A Graph-based Router for LLM Selections”
Tao Feng, Yanzhen Shen and Jiaxuan You · 2025
Closest in time.
“Fox-1 Technical Report”, 2025
Zijian Hu et al · 2025
Closest in time.
“Universal Model Routing for Efficient LLM Inference”, 2025
Wittawat Jitkrittum, Harikrishna Narasimhan, Ankit Rawat, Jeevesh Juneja, Zifeng Wang, Chen-Yu Lee, Pradeep Shenoy, Rina Panigrahy, Aditya Menon and Sanjiv Kumar · 2025
Closest in time.
“NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models”, 2025
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro and Wei Ping · 2025
Closest in time.
Seanie Lee, Dong Lee, Dominik Wagner, Minki Kang, Haebin Seong, Tobias Bocklet, Juho Lee and Sung Hwang · 2025
Closest in time.
“Making Text Embedders Few-Shot Learners”
Chaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Defu Lian, Yingxia Shao and Zheng Liu · 2025
Closest in time.
“LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing”, 2025
Yang Li · 2025
Closest in time.
“Generative Representational Instruction Tuning”
Niklas Muennighoff, Hongjin SU, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh and Douwe Kiela · 2025
Closest in time.
“RouteLLM: Learning to Route LLMs from Preference Data”
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph. Gonzalez, M Kadous and Ion Stoica · 2025
Closest in time.
“CARROT: A Cost Aware Rate Optimal Router”, 2025
Seamus Somerstep, Felipe Polo, Allysson de Oliveira, Prattyush Mangal, Mírian Silva, Onkar Bhardwaj, Mikhail Yurochkin and Subha Maity · 2025
Closest in time.
“MixLLM: Dynamic Routing in Mixed Large Language Models”, 2025
Xinyuan Wang, Yanchi Liu, Wei Cheng, Xujiang Zhao, Zhengzhang Chen, Wenchao Yu, Yanjie Fu and Haifeng Chen · 2025
Closest in time.
“Leveraging Uncertainty Estimation for Efficient LLM Routing”, 2025
Tuo Zhang, Asal Mehradfar, Dimitrios Dimitriadis and Salman Avestimehr · 2025
Closest in time.
“EmbedLLM: Learning Compact Representations of Large Language Models”
Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao and Kannan Ramchandran · 2025
Closest in time.