Fetching the paper…
Reading the bibliography…
Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Rapid object detection using a boosted cascade of simple features
P. Viola and M. Jones · 2001
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Supervised sequential classification under budget constraints
Kirill Trapeznikov and Venkatesh Saligrama · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung · 2016
Earlier work this paper cites.
Adaptive neural networks for fast test-time prediction
Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Multi-scale dense networks for resource efficient image classification
Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Weinberger · 2018
Earlier work this paper cites.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Earlier work this paper cites.
Predict responsibly: Improving fairness and accuracy by learning to defer
David Madras, Toniann Pitassi, and Richard Zemel · 2018
Earlier work this paper cites.
Learning to reweight examples for robust deep learning
Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun · 2018
Earlier work this paper cites.
Approximation algorithms for cascading prediction models
Matthew Streeter · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Earlier work this paper cites.
IDK cascades: Fast deep learning by learning not to overthink
Xin Wang, Yujia Luo, Daniel Crankshaw, Alexey Tumanov, Fisher Yu, and Joseph E. Gonzalez · 2018
Earlier work this paper cites.
SELFIE: Refurbishing unclean samples for robust deep learning
Hwanjun Song, Minseok Kim, and Jae-Gil Lee · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Consistent estimators for learning to defer to an expert
Hussein Mozannar and David Sontag · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Earlier work this paper cites.
The right tool for the job: Matching model and instance complexities
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith · 2020
Cited alongside, same era.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin · 2020
Cited alongside, same era.
Deep learning through the lens of example difficulty
Robert Baldock, Hartmut Maennel, and Behnam Neyshabur · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Cited alongside, same era.
When in doubt, summon the titans: Efficient inference with large models
Ankit Singh Rawat, Manzil Zaheer, Aditya Krishna Menon, Amr Ahmed, and Sanjiv Kumar · 2021
Cited alongside, same era.
FrugalGPT: How to use large language models while reducing cost and improving performance, 2023
Lingjiao Chen, Matei Zaharia, and James Zou · 2023
Later among the works it cites.
Matformer: Nested transformer for elastic inference, 2023
Devvrit, Sneha Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham Kakade, Ali Farhadi, and Prateek Jain · 2023
Later among the works it cites.
Tryage: Real-time, intelligent routing of user prompts to large language models, 2023
Surya Narayanan Hari and Matt Thomson · 2023
Later among the works it cites.
Towards anytime classification in early-exit architectures by enforcing conditional monotonicity
Metod Jazbec, James Urquhart Allingham, Dan Zhang, and Eric Nalisnick · 2023
Later among the works it cites.
When does confidence-based cascade deferral suffice?
Wittawat Jitkrittum, Neha Gupta, Aditya Krishna Menon, Harikrishna Narasimhan, Ankit Singh Rawat, and Sanjiv Kumar · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2021
Cited alongside, same era.
Estimating example difficulty using variance of gradients
Chirag Agarwal, Daniel D’souza, and Sara Hooker · 2022
Cited alongside, same era.
David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-dickstein, Kevin Murphy, and Charles Sutton · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Cited alongside, same era.
Babybear: Cheap inference triage for expensive language models, 2022
Leila Khalili, Yao You, and John Bohannon · 2022
Cited alongside, same era.
Findings of the 2022 conference on machine translation (wmt22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, et al · 2022
Cited alongside, same era.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al · 2022
Cited alongside, same era.
Efficient edge inference by selective query
Anil Kag, Igor Fedorov, Aditya Gangrade, Paul Whatmough, and Venkatesh Saligrama · 2023
Later among the works it cites.
Routing to the expert: Efficient reward-guided ensemble of large language models, 2023
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou · 2023
Later among the works it cites.
Federated automatic differentiation, 2023
Keith Rush, Zachary Charles, and Zachary Garrett · 2023
Later among the works it cites.
Large language model routing with benchmark datasets, 2023
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson, and Mikhail Yurochkin · 2023
Later among the works it cites.
Tabi: An efficient multi-level inference system for large language models
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo · 2023
Later among the works it cites.
Combating noisy labels with sample selection by mining high-discrepancy examples
Xiaobo Xia, Bo Han, Yibing Zhan, Jun Yu, Mingming Gong, Chen Gong, and Tongliang Liu · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem · 2024
Closest in time.
Hybrid LLM: Cost-efficient and quality-aware query routing
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah · 2024
Closest in time.
Minillm: Knowledge distillation of large language models, 2024
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2024
Closest in time.
Language model cascades: Token-level uncertainty and beyond
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, and Sanjiv Kumar · 2024
Closest in time.
Routerbench: A benchmark for multi-llm routing system, 2024
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay · 2024
Closest in time.
Orchestrallm: Efficient orchestration of language models for dialogue state tracking, 2024
Chia-Hsuan Lee, Hao Cheng, and Mari Ostendorf · 2024
Closest in time.
Theoretically grounded loss functions and algorithms for score-based multi-class abstention
Anqi Mao, Mehryar Mohri, and Yutao Zhong · 2024
Closest in time.
Jointly-learned exit and inference for a dynamic neural network
Florence Regol, Joud Chataoui, and Mark Coates · 2024
Closest in time.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Marija Sakota, Maxime Peyrard, and Robert West · 2024
Closest in time.
The bitter lesson
Rich Sutton · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, and others · 2024
Closest in time.
Sentence-level or token-level? a comprehensive study on knowledge distillation, 2024
Jingxuan Wei, Linzhuang Sun, Yichong Leng, Xu Tan, Bihui Yu, and Ruifeng Guo · 2024
Closest in time.
Large language model cascades with mixture of thought representations for cost-efficient reasoning
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao · 2024
Closest in time.