Fetching the paper…
Reading the bibliography…
We propose a new model for multi-token prediction in transformers, aiming to enhance sampling efficiency without compromising accuracy.
Foundations of the PARAFAC procedure: models and conditions for an explanatory multimodal factor analysis
Harshman, R · 1970
Earlier work this paper cites.
Tensor decompositions and applications
Kolda, T. G. and Bader, B. W · 2009
Earlier work this paper cites.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Earlier work this paper cites.
Tensor networks for dimensionality reduction and large-scale optimization: Part 1 low-rank tensor decompositions
Cichocki, A., Lee, N., Oseledets, I., Phan, A.-H., Zhao, Q., Mandic, D. P., et al · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Accelerating feedforward computation via parallel nonlinear equation solving
Song, Y., Meng, C., Liao, R., and Ermon, S · 2021
Earlier work this paper cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Dou, S., Zhou, E., Liu, Y., Gao, S., Zhao, J., Shen, W., Zhou, Y., Xi, Z., Wang, X., Fan, X., et al · 2023
Earlier work this paper cites.
Tinystories: How small can language models be and still speak coherent english?
Eldan, R. and Li, Y · 2023
Cited alongside, same era.
A practical survey on faster and lighter transformers
Fournier, Q., Caron, G. M., and Aloise, D · 2023
Cited alongside, same era.
Speed: Speculative pipelined execution for efficient decoding
Hooper, C., Kim, S., Mohammadzadeh, H., Genc, H., Keutzer, K., Gholami, A., and Shao, S · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Accelerating transformer inference for translation via parallel decoding
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E · 2023
Cited alongside, same era.
Break the sequential dependency of llm inference using lookahead decoding
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Closest in time.
Gao, X., Xie, W., Xiang, Y., and Ji, F · 2024
Closest in time.
Better & faster large language models via multi-token prediction
Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G · 2024
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2023
Cited alongside, same era.
Speculative streaming: Fast llm inference without auxiliary models
Bhendawade, N., Belousova, I., Fu, Q., Mason, H., Rastegari, M., and Najibi, M · 2024
Cited alongside, same era.
A survey on mixture of experts
Cai, W., Jiang, J., Wang, F., Tang, J., Kim, S., and Huang, J · 2024
Cited alongside, same era.
Layer skip: Enabling early exit inference and self-speculative decoding
Elhoushi, M., Shrivastava, A., Liskovich, D., Hosmer, B., Wasti, B., Lai, L., Mahmoud, A., Acun, B., Agarwal, S., Roman, A., et al · 2024
Cited alongside, same era.
A survey of text classification with transformers: How wide? how large? how long? how accurate? how expensive? how safe?
Fields, J., Chovanec, K., and Madiraju, P · 2024
Cited alongside, same era.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T
Cited in the paper.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L., Wolf, T., et al
Cited in the paper.
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., et al · 2024
Closest in time.
Dense training, sparse inference: Rethinking training of mixture-of-experts language models
Pan, B., Shen, Y., Liu, H., Mishra, M., Zhang, G., Oliva, A., Raffel, C., and Panda, R · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
OPT-Tree: Speculative decoding with adaptive draft tree structure
Wang, J., Su, Y., Li, J., Xia, Q., Ye, Z., Duan, X., Wang, Z., and Zhang, M · 2024
Closest in time.
Dyspec: Faster speculative decoding with dynamic token tree structure
Xiong, Y., Zhang, R., Li, Y., Wu, T., and Zou, L · 2024
Closest in time.