Fetching the paper…
Reading the bibliography…
Large language models (LLMs) excel in most NLP tasks but also require expensive cloud servers for deployment due to their size, while smaller models that can be deployed on lower cost (e.g., edge) devices, tend to lag behind in terms of response quality.
Optimal brain damage
Y. LeCun, J. Denker, and S. Solla · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
B. Hassibi, D. G. Stork, and G. J. Wolff · 1993
Earlier work this paper cites.
Improving the speed of neural networks on cpus
V. Vanhoucke, A. Senior, and M. Z. Mao · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, J. Dean, et al · 2015
Earlier work this paper cites.
Learning from imbalanced data: open challenges and future directions
B. Krawczyk · 2016
Earlier work this paper cites.
Do deep convolutional nets really need to be deep and convolutional?
G. Urban, K. J. Geras, S. E. Kahou, O. Aslan, S. Wang, R. Caruana, A. Mohamed, M. Philipose, and M. Richardson · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
B. Zoph and Q. V. Le · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko · 2018
Earlier work this paper cites.
Neural architecture search: A survey
T. Elsken, J. H. Metzen, and F. Hutter · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Fast structured decoding for sequence models
Z. Sun, Z. Li, H. Wang, D. He, Z. Lin, and Z. Deng · 2019
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
P. He, X. Liu, J. Gao, and W. Chen · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Binary cross entropy with deep learning technique for image classification
U. Ruby and V. Yendapalli · 2020
Cited alongside, same era.
Video object segmentation and tracking: A survey
R. Yao, G. Lin, S. Xia, J. Zhao, and Y. Zhou · 2020
Cited alongside, same era.
On the dangers of stochastic parrots : can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
W. Yuan, G. Neubig, and P. Liu · 2021
Cited alongside, same era.
A global analysis of metrics used for measuring performance in natural language processing
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun · 2022
Later among the works it cites.
Frugalgpt: How to use large language models while reducing cost and improving performance
L. Chen, M. Zaharia, and J. Zou · 2023
Later among the works it cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang, X. Ren, and B. Y. Lin · 2023
Later among the works it cites.
Speculative decoding with big little decoder
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Y. Leviathan, M. Kalman, and Y. Matias · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Blagec, G. Dorffner, M. Moradi, S. Ott, and M. Samwald · 2022
Cited alongside, same era.
LangChain, Oct. 2022
H. Chase · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Cited alongside, same era.
On efficient approximate queries over machine learning models
D. Ding, S. Amer-Yahia, and L. V. Lakshmanan · 2022
Cited alongside, same era.
The survey: Text generation models in deep learning
T. Iqbal and S. Qureshi · 2022
Cited alongside, same era.
Efficient edge inference by selective query
A. Kag, I. Fedorov, A. Gangrade, P. Whatmough, and V. Saligrama · 2022
Cited alongside, same era.
Scaling up trustless dnn inference with zero-knowledge proofs
D. Kang, T. Hashimoto, I. Stoica, and Y. Sun · 2022
Cited alongside, same era.
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Later among the works it cites.
Efficient deep learning: A survey on making deep learning models smaller, faster, and better
G. Menghani · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Gpt-4 is non-deterministic and moe is the likely reason why, Aug 2023
L. Skyward · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Efficient methods for natural language processing: A survey
M. Treviso, J.-U. Lee, T. Ji, B. v. Aken, Q. Cao, M. R. Ciosici, M. Hassid, K. Heafield, S. Hooker, C. Raffel, et al · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.