Fetching the paper…
Reading the bibliography…
We present EE-LLM, a framework for large-scale training and inference of early-exit large language models (LLMs).
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. C · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P. W., Zhang, Z., and Tu, Z · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Bengio, E., Bacon, P.-L., Pineau, J., and Precup, D · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Ba, J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Parikh, A., Täckström, O., Das, D., and Uszkoreit, J · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S., McDanel, B., and Kung, H. T · 2016
Earlier work this paper cites.
Structured attention networks
Kim, Y., Denton, C., Hoang, L., and Rush, A. M · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N. M., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Multi-scale dense networks for resource efficient image classification
Huang, G., Chen, D., Li, T., Wu, F., van der Maaten, L., and Weinberger, K. Q · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Earlier work this paper cites.
Shallow-deep networks: Understanding and mitigating network overthinking
Kaya, Y., Hong, S., and Dumitras, T · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A. P., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Pipedream: generalized pipeline parallelism for DNN training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Elf: An early-exiting framework for long-tailed classification
Duggal, R., Freitas, S., Dhamnani, S., Chau, D. H., and Sun, J · 2020
Cited alongside, same era.
Depth-adaptive transformer
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2020
Cited alongside, same era.
Dynabert: Dynamic BERT with adaptive width and depth
Hou, L., Huang, Z., Shang, L., Jiang, X., Chen, X., and Liu, Q · 2020
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments
Yuan, B., He, Y., Davis, J., Zhang, T., Dao, T., Chen, B., Liang, P., Ré, C., and Zhang, C · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M. T., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Later among the works it cites.
Semi-hfl: semi-supervised federated learning for heterogeneous devices
Zhong, Z., Wang, J., Bao, W., Zhou, J., Zhu, X., and Zhang, X · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fastbert: a self-distilling BERT with adaptive inference time
Liu, W., Zhou, P., Wang, Z., Zhao, Z., Deng, H., and Ju, Q · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Why should we add early exits to neural networks?
Scardapane, S., Scarpiniti, M., Baccarelli, E., and Uncini, A · 2020
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Deebert: Dynamic early exiting for accelerating BERT inference
Xin, J., Tang, R., Lee, J., Yu, Y., and Lin, J · 2020
Cited alongside, same era.
BERT loses patience: Fast and robust inference with early exit
Zhou, W., Xu, C., Ge, T., McAuley, J. J., Xu, K., and Wei, F · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K. A., Davis, J., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N. D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T. F., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M. S., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J. F., Ogut, G., Orr, L. J., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y. H., Ruiz, C., Ryan, J., R’e, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K. P., Tamkin, A., Taori, R., Thomas, A. W., Tramèr, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia, M. A., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P · 2021
Cited alongside, same era.
Later among the works it cites.
Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding
Bae, S., Ko, J., Song, H., and Yun, S.-Y · 2023
Closest in time.
Data-juicer: A one-stop data processing system for large language models
Chen, D., Huang, Y., Ma, Z., Chen, H., Pan, X., Ge, C., Gao, D., Xie, Y., Liu, Z., Gao, J., Li, Y., Ding, B., and Zhou, J · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2023
Closest in time.
Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference
Corro, L. D., Giorno, A. D., Agarwal, S., Yu, T., Awadallah, A. H., and Mukherjee, S · 2023
Closest in time.
Apparate: Rethinking early exits to tame latency-throughput tensions in ml serving
Dai, Y., Pan, R., Iyer, A., Li, K., and Netravali, R · 2023
Closest in time.
Refined open source dataset by data-juicer
Data-Juicer · 2023
Closest in time.
Jump to conclusions: Short-cutting transformers with linear transformations
Din, A. Y., Karidi, T., Choshen, L., and Geva, M · 2023
Closest in time.
The benefits of bad advice: Autocontrastive decoding across model layers
Gera, A., Friedman, R., Arviv, O., Gunasekara, C., Sznajder, B., Slonim, N., and Shnarch, E · 2023
Closest in time.
Smartbert: A promotion of dynamic early exiting mechanism for accelerating bert inference
Hu, B., Zhu, Y., Li, J., and Tang, S · 2023
Closest in time.
Scalefl: Resource-adaptive federated learning with heterogeneous clients
Ilhan, F., Su, G., and Liu, L · 2023
Closest in time.
Decoderlens: Layerwise interpretation of encoder-decoder transformers
Langedijk, A., Mohebbi, H., Sarti, G., Zuidema, W. H., and Jumelet, J · 2023
Closest in time.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., R’e, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L. J., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N. S., Guha, N., Chatterji, N. S., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T. F., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Pipefisher: Efficient training of large language models using pipelining and fisher information matrices
Osawa, K., Li, S., and Hoefler, T · 2023
Closest in time.
Accelerating transformer inference for translation via parallel decoding
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E · 2023
Closest in time.
Deed: Dynamic early exit on decoder for accelerating encoder-decoder transformer models
Tang, P., Zhu, P., Li, T., Appalaraju, S., Mahadevan, V., and Manmatha, R · 2023
Closest in time.
Internlm: A multilingual language model with progressively enhanced capabilities
Team, I · 2023
Closest in time.
Varshney, N., Chatterjee, A., Parmar, M., and Baral, C · 2023
Closest in time.
Hadskip: Homotopic and adaptive layer skipping of pre-trained language models for efficient inference
Wang, H., Wang, Y., Liu, T., Zhao, T., and Gao, J · 2023
Closest in time.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Closest in time.
A survey on dynamic neural networks for natural language processing
Xu, C. and McAuley, J. J · 2023
Closest in time.
Holmes: Towards distributed training across clusters with heterogeneous nic environment
Yang, F., Peng, S., Sun, N., Wang, F., Tan, K., Wu, F., Qiu, J., and Pan, A · 2023
Closest in time.
Learning to skip for language modeling
Zeng, D., Du, N., Wang, T., Xu, Y., Lei, T., Chen, Z., and Cui, C · 2023
Closest in time.
A simple romance between multi-exit vision transformer and token reduction
Liu, D., Kan, M., Shan, S., and Chen, X · 2024
Closest in time.