Fetching the paper…
Reading the bibliography…
Inference optimizations are critical for improving user experience and reducing infrastructure costs and power consumption.
Nothing clear enough to list yet.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference.”
Cited in the paper.
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Q. Tran, Y. Tay, and D. Metzler, “Confident adaptive language modeling,” in
Cited in the paper.
version: 2
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt, “Eliciting latent predictions from transformers with the tuned lens.”
Cited in the paper.
version: 1
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper, “Accelerating large language model decoding with speculative sampling.”
Cited in the paper.
version: 2
S. Kim, K. Mangalam, S. Moon, J. Canny, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer, “Big little transformer decoder.”
Cited in the paper.
J. Gante, “Assisted generation: a new direction toward low-latency text generation.”
Cited in the paper.
M. Stern, N. Shazeer, and J. Uszkoreit, “Blockwise parallel decoding for deep autoregressive models.”
Cited in the paper.
D. Lai and B. Lu, “Understanding autoregressive model for time series as a deterministic dynamic system,” in
Cited in the paper.
N. Shazeer, “Fast transformer decoding: One write-head is all you need.”
Cited in the paper.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…