Fetching the paper…
Reading the bibliography…
The computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment.
Remarks on Some Nonparametric Estimates of a Density Function
Rosenblatt, M · 1956
Earlier work this paper cites.
The correlation ratio as a new similarity measure for multimodal image registration
Roche, A., Malandain, G., Pennec, X., and Ayache, N · 1998
Earlier work this paper cites.
Bow-2000 datasheet
Graphcore · 2000
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2009
Earlier work this paper cites.
The unreasonable effectiveness of recurrent neural networks
Karpathy, A · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D · 2019
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Vig, J · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast neural network inference
Kurtz, M., Kopinsky, J., Gelashvili, R., Matveev, A., Carr, J., Goin, M., Leiserson, W., Moore, S., Nell, B., Shavit, N., and Alistarh, D · 2020
Cited alongside, same era.
O ( n ) O(n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
LM-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Closest in time.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
Iceformer: Accelerated inference with long-sequence transformers on CPUs
Mao, Y., Ester, M., and Li, K · 2023
Closest in time.
gpt-fast
Meta · 2023
Closest in time.
NVIDIA H100 datasheet
NVIDIA · 2023
Closest in time.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Cited alongside, same era.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Cited alongside, same era.
Combiner: Full attention transformer with sparse computation cost
Ren, H., Dai, H., Dai, Z., Yang, M., Leskovec, J., Schuurmans, D., and Dai, B · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Cited alongside, same era.
NVIDIA A10 datasheet
NVIDIA · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
Closest in time.
FlexGen: high-throughput generative inference of large language models with a single GPU
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Closest in time.
H 2 O: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Closest in time.
The needle in a haystack test
Dhinakaran, A · 2024
Closest in time.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Closest in time.
llama.cpp
Gerganov, G · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta AI · 2024
Closest in time.
ReLU strikes back: Exploiting activation sparsity in large language models
Mirzadeh, S. I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C., Tuzel, O., Samei, G., Rastegari, M., and Farajtabar, M · 2024
Closest in time.