Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have a natural role in answering complex queries about data streams, but the high computational cost of LLM inference makes them infeasible in many such tasks.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Universality versus cultural specificity of three emotion domains: Some evidence based on the cascading model of emotional intelligence
Shao, B., Doucet, L., and Caruso, D. R · 2015
Earlier work this paper cites.
Hate Speech Dataset from a White Supremacy Forum
de Gibert, O., Perez, N., García-Pablos, A., and Cuadros, M · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A · 2018
Earlier work this paper cites.
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Earlier work this paper cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Li, L., Lin, Y., Chen, D., Ren, S., Li, P., Zhou, J., and Sun, X · 2020
Cited alongside, same era.
Fastbert: a self-distilling bert with adaptive inference time
Liu, W., Zhou, P., Wang, Z., Zhao, Z., Deng, H., and Ju, Q · 2020
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Deebert: Dynamic early exiting for accelerating bert inference
Xin, J., Tang, R., Lee, J., Yu, Y., and Lin, J · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Frugalgpt: How to use large language models while reducing cost and improving performance
Chen, L., Zaharia, M., and Zou, J · 2023
Later among the works it cites.
Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2023
Later among the works it cites.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Later among the works it cites.
When does confidence-based cascade deferral suffice?
Jitkrittum, W., Gupta, N., Menon, A. K., Narasimhan, H., Rawat, A. S., and Kumar, S · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online learning: A comprehensive survey
Hoi, S. C., Sahoo, D., Lu, J., and Zhao, P · 2021
Cited alongside, same era.
When in doubt, summon the titans: Efficient inference with large models
Rawat, A. S., Zaheer, M., Menon, A. K., Ahmed, A., and Kumar, S · 2021
Cited alongside, same era.
Babybear: Cheap inference triage for expensive language models
Khalili, L., You, Y., and Bohannon, J · 2022
Cited alongside, same era.
Model cascading: Towards jointly improving efficiency and accuracy of NLP systems
Varshney, N. and Baral, C · 2022
Cited alongside, same era.
Wisdom of committees: An overlooked approach to faster and more accurate models
Wang, X., Kondratyuk, D., Christiansen, E., Kitani, K. M., Movshovitz-Attias, Y., and Eban, E · 2022
Cited alongside, same era.
Lifting the curse of capacity gap in distilling large language models
Zhang, C., Yang, Y., Liu, J., Wang, J., Wu, W., Wang, B., and Song, D · 2022
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al
Cited in the paper.
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
Liu, X., Hu, L., Bailis, P., Stoica, I., Deng, Z., Cheung, A., and Zhang, H · 2023
Later among the works it cites.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
Cache & distil: Optimising api calls to large language models
Ramírez, G., Lindemann, M., Birch, A., and Titov, I · 2023
Later among the works it cites.
Stogiannidis, I., Vassos, S., Malakasiotis, P., and Androutsopoulos, I · 2023
Later among the works it cites.
Small models are valuable plug-ins for large language models
Xu, C., Xu, Y., Wang, S., Liu, Y., Zhu, C., and McAuley, J · 2023
Later among the works it cites.
Language model cascades: Token-level uncertainty and beyond
Gupta, N., Narasimhan, H., Jitkrittum, W., Rawat, A. S., Menon, A. K., and Kumar, S · 2024
Closest in time.