Fetching the paper…
Reading the bibliography…
We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the generation of multiple tokens from each transformer call.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 1911
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
M. Stern, N. Shazeer, and J. Uszkoreit · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2020
Cited alongside, same era.
The depth-to-width interplay in self-attention
Y. Levine, N. Wies, O. Sharir, H. Bata, and A. Shashua · 2020
Cited alongside, same era.
Predictive sampling with forecasting autoregressive models
A. Wiggers and E. Hoogeboom · 2020
Cited alongside, same era.
Vivit: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lucic, and C. Schmid · 2021
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Later among the works it cites.
Lossless acceleration for seq2seq generation with aggressive decoding
T. Ge, H. Xia, X. Sun, S. Chen, and F. Wei · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Cited alongside, same era.
Accelerating feedforward computation via parallel nonlinear equation solving
Y. Song, C. Meng, R. Liao, and S. Ermon · 2021
Cited alongside, same era.
Y. Leviathan, M. Kalman, and Y. Matias · 2022
Later among the works it cites.
Efficiently scaling transformer inference
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He · 2022
Later among the works it cites.