Fetching the paper…
Reading the bibliography…
To reduce the computation cost and the energy consumption in large language models (LLM), skimming-based acceleration dynamically drops unimportant tokens of the input sequence progressively along layers of the LLM while preserving the tokens of semantic importance.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
Liu, W.; Zhou, P.; Zhao, Z.; Wang, Z.; Deng, H.; and Ju, Q. 2020 · 2004
Earlier work this paper cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Xin, J.; Tang, R.; Lee, J.; Yu, Y.; and Lin, J. 2020 · 2004
Earlier work this paper cites.
Remote timing attacks are practical
Brumley, D.; and Boneh, D. 2005 · 2005
Earlier work this paper cites.
A Panda? No, It’s a Sloth: Slowdown Attacks on Adaptive Multi-Exit Neural Network Inference
Hong, S.; Kaya, Y.; Modoranu, I.-V.; and Dumitraş, T. 2020 · 2010
Earlier work this paper cites.
Length-adaptive transformer: Train once with length drop, use anytime with search
Kim, G.; and Cho, K. 2020 · 2010
Earlier work this paper cites.
Cache attacks enable bulk key recovery on the cloud
Inci, M. S.; Gulmezoglu, B.; Irazoqui, G.; Eisenbarth, T.; and Sunar, B. 2016 · 2016
Earlier work this paper cites.
Skip rnn: Learning to skip state updates in recurrent neural networks
Campos, V.; Jou, B.; Giró-i Nieto, X.; Torres, J.; and Chang, S.-F. 2017 · 2017
Earlier work this paper cites.
Spatially adaptive computation time for residual networks
Figurnov, M.; Collins, M. D.; Zhu, Y.; Zhang, L.; Huang, J.; Vetrov, D.; and Salakhutdinov, R. 2017 · 2017
Earlier work this paper cites.
Multi-scale dense networks for resource efficient image classification
Huang, G.; Chen, D.; Li, T.; Wu, F.; Van Der Maaten, L.; and Weinberger, K. Q. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Yu, A. W.; Lee, H.; and Le, Q. V. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Speed reading: Learning to read forbackward via shuttle
Fu, T.-J.; and Ma, W.-Y. 2018 · 2018
Cited alongside, same era.
PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination
Goyal, S.; Choudhury, A. R.; Raje, S.; Chakaravarthy, V.; Sabharwal, Y.; and Verma, A. 2020 · 2020
Later among the works it cites.
Ilfo: Adversarial attack on adaptive neural networks
Haque, M.; Chauhan, A.; Liu, C.; and Yang, W. 2020 · 2020
Later among the works it cites.
Bert loses patience: Fast and robust inference with early exit
Zhou, W.; Xu, C.; Ge, T.; McAuley, J.; Xu, K.; and Wei, F. 2020 · 2020
Later among the works it cites.
Tr-bert: Dynamic token reduction for accelerating bert inference
Ye, D.; Lin, Y.; Huang, Y.; and Sun, M. 2021 · 2021
Later among the works it cites.
Transkimmer: Transformer learns to layer-wise skim
Guan, Y.; Li, Z.; Leng, J.; Lin, Z.; and Guo, M. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Cited alongside, same era.
Fast and accurate text classification: Skimming, rereading and early stopping
Yu, K.; Liu, Y.; Schwing, A. G.; and Peng, J. 2018 · 2018
Cited alongside, same era.
Shallow-deep networks: Understanding and mitigating network overthinking
Kaya, Y.; Hong, S.; and Dumitras, T. 2019 · 2019
Cited alongside, same era.
Non-profiled deep learning-based side-channel attacks with sensitivity analysis
Timon, B. 2019 · 2019
Cited alongside, same era.
NMTSloth: understanding and testing efficiency degradation of neural machine translation systems
Chen, S.; Liu, C.; Haque, M.; Song, Z.; and Yang, W. 2022a
Cited in the paper.
NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation Models
Chen, S.; Song, Z.; Haque, M.; Liu, C.; and Yang, W. 2022b
Cited in the paper.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018a
Cited in the paper.
Haque, M.; Yadlapalli, Y.; Yang, W.; and Liu, C. 2022 · 2022
Later among the works it cites.
Learned token pruning for transformers
Kim, S.; Shen, S.; Thorsley, D.; Gholami, A.; Kwon, W.; Hassoun, J.; and Keutzer, K. 2022 · 2022
Later among the works it cites.
SlowBERT: Slow-down Attacks on Input-adaptive Multi-exit BERT
Zhang, S.; Pan, X.; Zhang, M.; and Yang, M. 2023 · 2023
Closest in time.