Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are expensive to deploy.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 1911
Earlier work this paper cites.
The truncated svd as a method for regularization
P. C. Hansen · 1987
Earlier work this paper cites.
GLU variants improve transformer
N. Shazeer · 2002
Earlier work this paper cites.
Understanding deep architectures using a recursive convolutional network
D. Eigen, J. T. Rolfe, R. Fergus, and Y. LeCun · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
S. Han, J. Pool, J. Tran, and W. J. Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. G. Howard, H. Adam, and D. Kalenichenko · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network
A. Sherstinsky · 2018
Earlier work this paper cites.
Deep equilibrium models
S. Bai, J. Z. Kolter, and V. Koltun · 2019
Earlier work this paper cites.
Recurrent stacking of layers for compact neural machine translation models
R. Dabre and A. Fujita · 2019
Earlier work this paper cites.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Earlier work this paper cites.
Dynamic recursive neural network
Q. Guo, Z. Yu, Y. Wu, D. Liang, H. Qin, and J. Yan · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap · 2019
Earlier work this paper cites.
Learning implicitly recurrent cnns through parameter sharing
P. Savarese and M. Maire · 2019
Earlier work this paper cites.
Tied transformers: Neural machine translation with shared encoder and decoder
Y. Xia, T. He, X. Tan, F. Tian, D. He, and T. Qin · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Earlier work this paper cites.
Depth-adaptive transformer
M. Elbayad, J. Gu, E. Grave, and M. Auli · 2020
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
What’s hidden in a randomly weighted neural network?
V. Ramanujan, M. Wortsman, A. Kembhavi, A. Farhadi, and M. Rastegari · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2020
Cited alongside, same era.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
X. Liu, K. Ji, Y. Fu, Z. Du, Z. Yang, and J. Tang · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher, 2021
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. Hechtman, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving · 2021
Customizable combination of parameter-efficient modules for multi-task learning
H. Wang, T. Sun, C. Jin, Y. Wang, Y. Fan, Y. Xu, Y. Du, and C. Fan · 2023
Later among the works it cites.
f-divergence minimization for sequence-level knowledge distillation
Y. Wen, Z. Li, W. Du, and L. Mou · 2023
Later among the works it cites.
Learning to skip for language modeling
D. Zeng, N. Du, T. Wang, Y. Xu, T. Lei, Z. Chen, and C. Cui · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem · 2024
Closest in time.
Reducing transformer key-value cache size with cross-layer attention
W. Brandon, M. Mishra, A. Nrusimha, R. Panda, and J. Ragan-Kelley · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Consistent accelerated inference via confident adaptive transformers
T. Schuster, A. Fisch, T. Jaakkola, and R. Barzilay · 2021
Cited alongside, same era.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
A. Schwarzschild, E. Borgnia, A. Gupta, F. Huang, U. Vishkin, M. Goldblum, and T. Goldstein · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Cited alongside, same era.
Edgeformer: A parameter-efficient transformer for on-device seq2seq generation
T. Ge, S. Chen, and F. Wei · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
[re] end-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking
S. M. McLeish and L. Tran-Thanh · 2022
Cited alongside, same era.
EE-LLM: large-scale training and inference of early-exit large language models with 3d parallelism
Y. Chen, X. Pan, Y. Li, B. Ding, and J. Zhou · 2024
Closest in time.
Compressed chain of thought: Efficient reasoning through dense representations
J. Cheng and B. Van Durme · 2024
Closest in time.
Moeut: Mixture-of-experts universal transformers
R. Csordás, K. Irie, J. Schmidhuber, C. Potts, and C. D. Manning · 2024
Closest in time.
The llama 3 herd of models
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Rozière, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. M. Kloumann, I. Misra, I. Evtimov, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, and et al · 2024
Closest in time.
Layerskip: Enabling early exit inference and self-speculative decoding
M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, A. A. Aly, B. Chen, and C. Wu · 2024
Closest in time.
Looped transformers for length generalization
Y. Fan, Y. Du, K. Ramchandran, and K. Lee · 2024
Closest in time.
Mixture-of-loras: An efficient multitask tuning method for large language models
W. Feng, C. Hao, Y. Zhang, Y. Han, and H. Wang · 2024
Closest in time.
Lazyllm: Dynamic token pruning for efficient long context LLM inference
Q. Fu, M. Cho, T. Merth, S. Mehta, M. Rastegari, and M. Najibi · 2024
Closest in time.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?, 2024
K. Gatmiry, N. Saunshi, S. J. Reddi, S. Jegelka, and S. Kumar · 2024
Closest in time.
Zamba: A compact 7b SSM hybrid model
P. Glorioso, Q. Anthony, Y. Tokpanov, J. Whittington, J. Pilault, A. Ibrahim, and B. Millidge · 2024
Closest in time.
Think before you speak: Training language models with pause tokens
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Closest in time.
Minillm: Knowledge distillation of large language models
Y. Gu, L. Dong, F. Wei, and M. Huang · 2024
Closest in time.
Training large language models to reason in a continuous latent space
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Closest in time.
SOLAR 10.7b: Scaling large language models with simple yet effective depth up-scaling
S. Kim, D. Kim, C. Park, W. Lee, W. Song, Y. Kim, H. Kim, Y. Kim, H. Lee, J. Kim, C. Ahn, S. Yang, S. Lee, H. Park, G. Gim, M. Cha, H. Lee, and S. Kim · 2024
Closest in time.
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Z. Liu, C. Zhao, F. N. Iandola, C. Lai, Y. Tian, I. Fedorov, Y. Xiong, E. Chang, Y. Shi, R. Krishnamoorthi, L. Lai, and V. Chandra · 2024
Closest in time.
Pissa: Principal singular values and singular vectors adaptation of large language models
F. Meng, Z. Wang, and M. Zhang · 2024
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Closest in time.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
D. Raposo, S. Ritter, B. A. Richards, T. P. Lillicrap, P. C. Humphreys, and A. Santoro · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. P. Lillicrap, J. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, I. Antonoglou, R. Anil, S. Borgeaud, A. M. Dai, K. Millican, E. Dyer, M. Glaese, T. Sottiaux, B. Lee, F. Viola, M. Reynolds, Y. Xu, J. Molloy, J. Chen, M. Isard, P. Barham, T. Hennigan, R. McIlroy, M. Johnson, J. Schalkwyk, E. Collins, E. Rutherford, E. Moreira, K. Ayoub, M. Goel, C. Meyer, G. Thornton, Z. Yang, H. Michalewski, Z. Abbas, N. Schucher, A. Anand, R. Ives, J. Keeling, K. Lenc, S. Haykal, S. Shakeri, P. Shyam, A. Chowdhery, R. Ring, S. Spencer, E. Sezener, and et al · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
M. Rivière, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev, M. Hoffman, S. Thakoor, J. Grill, B. Neyshabur, O. Bachem, A. Walton, A. Severyn, A. Parrish, A. Ahmad, A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson, B. Bastian, B. Piot, B. Wu, B. Royal, C. Chen, C. Kumar, C. Perry, C. Welty, C. A. Choquette-Choo, D. Sinopalnikov, D. Weinberger, D. Vijaykumar, D. Rogozinska, D. Herbison, E. Bandy, E. Wang, E. Noland, E. Moreira, E. Senter, E. Eltyshev, F. Visin, G. Rasskin, G. Wei, G. Cameron, G. Martins, H. Hashemi, H. Klimczak-Plucinska, H. Batra, H. Dhand, I. Nardini, J. Mein, J. Zhou, J. Svensson, J. Stanway, J. Chan, J. P. Zhou, J. Carrasqueira, J. Iljazi, J. Becker, J. Fernandez, J. van Amersfoort, J. Gordon, J. Lipschultz, J. Newlan, J. Ji, K. Mohamed, K. Badola, K. Black, K. Millican, K. McDonell, K. Nguyen, K. Sodhia, K. Greene, L. L. Sjösund, L. Usui, L. Sifre, L. Heuermann, L. Lago, and L. McNealus · 2024
Closest in time.
On the inductive bias of stacking towards improving reasoning, 2024
N. Saunshi, S. Karp, S. Krishnan, S. Miryoosefi, S. J. Reddi, and S. Kumar · 2024
Closest in time.
Leveraging adapter for parameter-efficient asr encoder
K. Shim, J. Lee, and H. Kim · 2024
Closest in time.
You only cache once: Decoder-decoder architectures for language models
Y. Sun, L. Dong, Y. Zhu, S. Huang, W. Wang, S. Ma, Q. Zhang, J. Wang, and F. Wei · 2024
Closest in time.
DEED: dynamic early exit on decoder for accelerating encoder-decoder transformer models
P. Tang, P. Zhu, T. Li, S. Appalaraju, V. Mahadevan, and R. Manmatha · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Efficient large language models: A survey
Z. Wan, X. Wang, C. Liu, S. Alam, Y. Zheng, J. Liu, Z. Qu, S. Yan, Y. Zhu, Q. Zhang, M. Chowdhury, and M. Zhang · 2024
Closest in time.
Looped transformers are better at learning learning algorithms
L. Yang, K. Lee, R. D. Nowak, and D. Papailiopoulos · 2024
Closest in time.
Draft& verify: Lossless large language model acceleration via self-speculative decoding
J. Zhang, J. Wang, H. Li, L. Shou, K. Chen, G. Chen, and S. Mehrotra · 2024
Closest in time.
A survey on efficient inference for large language models
Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y. Lou, L. Wang, Z. Yuan, X. Li, S. Yan, G. Dai, X. Zhang, Y. Dong, and Y. Wang · 2024
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein · 2025
Closest in time.