Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved great success in solving difficult tasks across many domains, but such success comes with a high computation cost, and inference latency.
Adabert: Task-adaptive bert compression with differentiable neural architecture search
Chen, D., Li, Y., Qiu, M., Wang, Z., Li, B., Ding, B., Deng, H., Huang, J., Lin, W., and Zhou, J · 2001
Earlier work this paper cites.
Trade: Transformers for density estimation
Fakoor, R., Chaudhari, P., Mueller, J., and Smola, A. J · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, B. and Brockett, C · 2005
Earlier work this paper cites.
Hat: Hardware-aware transformers for efficient natural language processing
Wang, H., Wu, Z., Liu, Z., Cai, H., Zhu, L., Gan, C., and Han, S · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2006
Earlier work this paper cites.
Convex and semi-nonnegative matrix factorizations
Ding, C. H., Li, T., and Jordan, M. I · 2008
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Earlier work this paper cites.
First quora dataset release: question pairs (2017)
Shankar, I., Nikhil, D., and Kornel, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2017
Earlier work this paper cites.
Data-dependent coresets for compressing neural networks with applications to generalization bounds
Baykal, C., Liebenwein, L., Gilitschenski, I., Feldman, D., and Rus, D · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Kernelized convex hull approximation and its applications in data description tasks
Huang, C., Wu, Y., Min, G., and Ying, Y · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
Neuroinspired unsupervised learning and pruning with subquantum cbram arrays
Shi, Y., Nguyen, L., Oh, S., Liu, X., Koushan, F., Jameson, J. R., and Kuzum, D · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
Provable filter pruning for efficient neural networks
Liebenwein, L., Baykal, C., Lang, H., Feldman, D., and Rus, D · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Data-independent neural pruning via coresets
Mussay, B., Osadchy, M., Braverman, V., Zhou, S., and Feldman, D · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
The evolved transformer
So, D., Le, Q., and Liang, C · 2019
Cited alongside, same era.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Later among the works it cites.
Bert loses patience: Fast and robust inference with early exit
Zhou, W., Xu, C., Ge, T., McAuley, J., Xu, K., and Wei, F · 2020
Later among the works it cites.
Unsupervised pulsenet: Automated pruning of convolutional neural networks by k-means clustering
Browne, D., Giering, M., and Prestwich, S · 2021
Later among the works it cites.
Chasing sparsity in vision transformers: An end-to-end exploration
Chen, T., Cheng, Y., Gan, Z., Yuan, L., Zhang, L., and Wang, Z · 2021
Later among the works it cites.
Compressing large-scale transformer-based models: A case study on bert
Ganesh, P., Chen, Y., Lou, X., Khan, M. A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., and Winslett, M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Cited alongside, same era.
Structured pruning of large language models
Wang, Z., Wohlwend, J., and Lei, T · 2019
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Cited alongside, same era.
Fast convex pruning of deep neural networks
Aghasi, A., Abdi, A., and Romberg, J · 2020
Cited alongside, same era.
Pulsenetone: fast unsupervised pruning of convolutional neural networks for remote sensing
Browne, D., Giering, M., and Prestwich, S · 2020
Cited alongside, same era.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Ham, T. J., Lee, Y., Seo, S. H., Kim, S., Choi, H., Jung, S. J., and Lee, J. W · 2021
Later among the works it cites.
Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks
Hubara, I., Chmiel, B., Island, M., Banner, R., Naor, J., and Soudry, D · 2021
Later among the works it cites.
Retrieve: Coreset selection for efficient and robust semi-supervised learning
Killamsetty, K., Zhao, X., Chen, F., and Iyer, R · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Later among the works it cites.
Ebert: Efficient bert inference with dynamic structured pruning
Liu, Z., Li, F., Li, G., and Cheng, J · 2021
Later among the works it cites.
Data-independent structured pruning of neural networks via coresets
Mussay, B., Feldman, D., Zhou, S., Braverman, V., and Osadchy, M · 2021
Later among the works it cites.
Searching for efficient transformers for language modeling
So, D., Mańke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V · 2021
Later among the works it cites.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Tambe, T., Hooper, C., Pentecost, L., Jia, T., Yang, E.-Y., Donato, M., Sanh, V., Whatmough, P., Rush, A. M., Brooks, D., et al · 2021
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2021
Later among the works it cites.
Nas-bert: task-agnostic and adaptive-size bert compression with neural architecture search
Xu, J., Tan, X., Luo, R., Song, K., Li, J., Qin, T., and Liu, T.-Y · 2021
Later among the works it cites.
Mlpruning: A multilevel structured pruning framework for transformer-based models
Yao, Z., Ma, L., Shen, S., Keutzer, K., and Mahoney, M. W · 2021
Later among the works it cites.
Autotinybert: Automatic hyper-parameter optimization for efficient pre-trained language models
Yin, Y., Chen, C., Shang, L., Jiang, X., Chen, X., and Liu, Q · 2021
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Later among the works it cites.
Heat: Hardware-efficient automatic tensor decomposition for transformer compression
Gu, J., Keller, B., Kossaifi, J., Anandkumar, A., Khailany, B., and Pan, D. Z · 2022
Later among the works it cites.
Tackling provably hard representative selection via graph neural networks
Kazemi, S. M., Tsitsulin, A., Esfandiari, H., Bateni, M., Ramachandran, D., Perozzi, B., and Mirrokni, V · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers
Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A · 2022
Later among the works it cites.
Littlebird: Efficient faster & longer transformer for question answering
Lee, M., Han, K., and Shin, M. C · 2022
Later among the works it cites.
Greedy-layer pruning: Speeding up transformer models for natural language processing
Peer, D., Stabinger, S., Engl, S., and Rodríguez-Sánchez, A · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Xia, M., Zhong, Z., and Chen, D · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance
Zhang, Q., Zuo, S., Liang, C., Bukharin, A., He, P., Chen, W., and Zhao, T · 2022
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Sajjad, H., Dalvi, F., Durrani, N., and Nakov, P · 2023
Closest in time.