Fetching the paper…
Reading the bibliography…
Inference from large autoregressive models like Transformers is slow - decoding K tokens takes K serial runs of the model.
Speculative computation, parallelism, and functional programming
Burton, F. W · 1985
Earlier work this paper cites.
Computer Architecture: A Quantitative Approach
Hennessy, J. L. and Patterson, D. A · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P. T., and Robinson, T · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N. M · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Cited alongside, same era.
Controlling computation versus quality for neural sequence models
Bapna, A., Arivazhagan, N., and Firat, O · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Sparse is enough in scaling transformers
Jaszczur, S., Chowdhery, A., Mohiuddin, A., Kaiser, L., Gajewski, W., Michalewski, H., and Kanerva, J · 2021
Later among the works it cites.
Consistent accelerated inference via confident adaptive transformers
Schuster, T., Fisch, A., Jaakkola, T., and Barzilay, R · 2021
Later among the works it cites.
Primer: Searching for efficient transformers for language modeling
So, D. R., Ma’nke, W., Liu, H., Dai, Z., Shazeer, N. M., and Le, Q. V · 2021
Later among the works it cites.
Instantaneous grammatical error correction with shallow aggressive decoding
Sun, X., Ge, T., Wei, F., and Wang, H · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N. M., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B. C., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., García, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Díaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K. S., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why should we add early exits to neural networks?
Scardapane, S., Scarpiniti, M., Baccarelli, E., and Uncini, A · 2020
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Dehghani, M., Arnab, A., Beyer, L., Vaswani, A., and Tay, Y · 2021
Cited alongside, same era.
Dynamic neural networks: A survey
Han, Y., Huang, G., Song, S., Yang, L., Wang, H., and Wang, Y · 2021
Cited alongside, same era.
Closest in time.
Scaling up models and data with t5x and seqio
Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., García, X., Ni, J., Chen, A., Kenealy, K., Clark, J., Lee, S., Garrette, D. H., Lee-Thorp, J., Raffel, C., Shazeer, N. M., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J. B., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A · 2022
Closest in time.
Lamda: Language models for dialog applications
Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N. M., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C.-C., Krivokon, I. A., Rusch, W. J., Pickett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Søraker, J. H., Zevenbergen, B., Prabhakaran, V., Díaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V. O., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., hsin Chi, E. H., and Le, Q · 2022
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B. C., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Closest in time.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. M · 2023
Closest in time.