Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) and other large foundation models have achieved noteworthy success, but their size exacerbates existing resource consumption and latency challenges.
Cache memories
A. J. Smith · 1982
Earlier work this paper cites.
A survey of web caching schemes for the internet
J. Wang · 1999
Earlier work this paper cites.
Evaluating content management techniques for web proxy caches
M. Arlitt, L. Cherkasova, J. Dilley, R. Friedrich, and T. Jin · 2000
Earlier work this paper cites.
Popularity-aware greedy dual-size web proxy caching algorithms
S. Jin and A. Bestavros · 2000
Earlier work this paper cites.
LRFU: A spectrum of policies that subsumes the least recently used and least frequently used policies
D. Lee, J. Choi, J.-H. Kim, S. H. Noh, S. L. Min, Y. Cho, and C. S. Kim · 2001
Earlier work this paper cites.
Web cache management based on the expected cost of web objects
H. Bahn · 2005
Earlier work this paper cites.
Operating Systems: Internals and Design Principles , volume 9
W. Stallings and G. K. Paul · 2012
Earlier work this paper cites.
High dimensional statistics
P. Rigollet and J.-C. Hütter · 2015
Earlier work this paper cites.
Semantic search on text and knowledge bases
H. Bast, B. Buchhold, E. Haussmann, et al · 2016
Earlier work this paper cites.
An overview of modern cache memory and performance analysis of replacement policies
S. Kumar and P. Singh · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Deep-reinforcement-learning-based optimization for cache-enabled opportunistic interference alignment wireless networks
Y. He, Z. Zhang, F. R. Yu, N. Zhao, H. Yin, V. C. Leung, and Y. Zhang · 2017
Earlier work this paper cites.
Learn to cache: Machine learning for network edge caching in the big data era
Z. Chang, L. Lei, Z. Zhou, S. Mao, and T. Ristaniemi · 2018
Earlier work this paper cites.
Multi-agent reinforcement learning for efficient content caching in mobile d2d networks
W. Jiang, G. Feng, S. Qin, T. S. P. Yum, and G. Cao · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Pre-training tasks for embedding-based large-scale retrieval, 2020
W.-C. Chang, F. X. Yu, Y.-W. Chang, Y. Yang, and S. Kumar · 2020
Earlier work this paper cites.
General purpose text embeddings from pre-trained language models for scalable inference, 2020
J. Du, M. Ott, H. Li, X. Zhou, and V. Stoyanov · 2020
Cited alongside, same era.
The cost of training NLP models: A concise overview
O. Sharir, B. Peleg, and Y. Shoham · 2020
Cited alongside, same era.
Learning to cache and caching to learn: Regret analysis of caching algorithms
A. Bura, D. Rengarajan, D. Kalathil, S. Shakkottai, and J.-F. Chamberland · 2021
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference, 2021
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2021
Cited alongside, same era.
Online caching with optimal switching regret
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
StackLLaMA: An RL fine-tuned LLaMA model for Stack Exchange question and answering, 2023
E. Beeching, Y. Belkada, K. Rasul, L. Tunstall, L. von Werra, N. Rajani, and N. Lambert · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Regret-optimal online caching for adversarial and stochastic arrivals
F. Z. Faizal, P. Singh, N. Karamchandani, and S. Moharir · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Mukhopadhyay and A. Sinha · 2021
Cited alongside, same era.
Carbon emissions and large neural network training
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Cited alongside, same era.
Applying machine learning techniques for caching in next-generation edge networks: A comprehensive survey
J. Shuja, K. Bilal, W. Alasmary, H. Sinky, and E. Alanazi · 2021
Cited alongside, same era.
Single-layer vision transformers for more accurate early exits with less overhead, 2022
A. Bakhtiarnia, Q. Zhang, and A. Iosifidis · 2022
Cited alongside, same era.
On the opportunities and risks of foundation models, 2022
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang · 2022
Cited alongside, same era.
PaLM: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
Closest in time.
GPTQ: Accurate post-training quantization for generative pre-trained transformers, 2023
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2023
Closest in time.
PaLM-2 technical report
Google · 2023
Closest in time.
Evaluating embedding APIs for information retrieval, 2023
E. Kamalloo, X. Zhang, O. Ogundepo, N. Thakur, D. Alfonso-Hermelo, M. Rezagholizadeh, and J. Lin · 2023
Closest in time.
Big little transformer decoder, 2023
S. Kim, K. Mangalam, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer · 2023
Closest in time.
Openassistant conversations–democratizing large language model alignment
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, et al · 2023
Closest in time.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Closest in time.
Capabilities of GPT-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Embedding recycling for language models, 2023
J. Saad-Falcon, A. Singh, L. Soldaini, M. D’Arcy, A. Cohan, and D. Downey · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT, 2023
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, H. Peng, J. Li, J. Wu, Z. Liu, P. Xie, C. Xiong, J. Pei, P. S. Yu, and L. Sun · 2023
Closest in time.