Fetching the paper…
Reading the bibliography…
Autoregressive decoding strategy is a commonly used method for text generation tasks with pre-trained language models, while early-exiting is an effective approach to speedup the inference stage.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. pp. 311–318 (Jul 2002). https://doi.org/10.3115/1073083.1073135
2002
Earlier work this paper cites.
Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Proceedings of the 42nd annual meeting of the association for computational linguistics. pp. 74–81 (Jul 2004)
2004
Earlier work this paper cites.
Hinton, G., Vinyals, O., Dean, J.: Distilling the knowledge in a neural network (Mar 2015)
2015
Earlier work this paper cites.
Srinivas, S., Babu, R.V.: Data-free parameter pruning for deep neural networks. In: British Machine Vision Conference (2015)
2015
Earlier work this paper cites.
Han, S., Mao, H., Dally, W.J.: Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. In: International Conference on Learning Representations (2016)
2016
Earlier work this paper cites.
Nallapati, R., Zhou, B., dos Santos, C., Gulcehre, C., Xiang, B.: Abstractive text summarization using sequence-to-sequence RNNs and beyond. In: Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning. pp. 280–290 (Aug 2016). https://doi.org/10.18653/v1/K16-1028
2016
Earlier work this paper cites.
Ranzato, M., Chopra, S., Auli, M., Zaremba, W.: Sequence level training with recurrent neural networks. In: International Conference on Learning Representations (2016)
2016
Earlier work this paper cites.
Teerapittayanon, S., McDanel, B., Kung, H.: Branchynet: Fast inference via early exiting from deep neural networks. In: 2016 23rd International Conference on Pattern Recognition (ICPR). pp. 2464–2469 (2016). https://doi.org/10.1109/ICPR.2016.7900006
2016
Earlier work this paper cites.
Bolukbasi, T., Wang, J., Dekel, O., Saligrama, V.: Adaptive neural networks for efficient inference. In: Proceedings of the 34th International Conference on Machine Learning. p. 527–536 (2017)
2017
Earlier work this paper cites.
Li, H., Kadav, A., Durdanovic, I., Samet, H., Graf, H.P.: Pruning filters for efficient convnets. In: International Conference on Learning Representations (2017)
2017
Earlier work this paper cites.
Novikova, J., Dušek, O., Rieser, V.: The E2E dataset: New challenges for end-to-end generation. In: Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. pp. 201–206 (Aug 2017). https://doi.org/10.18653/v1/W17-5525
2017
Earlier work this paper cites.
Narayan, S., Cohen, S.B., Lapata, M.: Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. pp. 1797–1807 (Oct-Nov 2018). https://doi.org/10.18653/v1/D18-1206
2018
Earlier work this paper cites.
Radford, A., Narasimhan, K.: Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
Wang, X., Yu, F., Dou, Z.Y., Darrell, T., Gonzalez, J.E.: Skipnet: Learning dynamic routing in convolutional networks. In: The European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 4171–4186 (Jun 2019). https://doi.org/10.18653/v1/N19-1423
2019
Earlier work this paper cites.
Kaya, Y., Hong, S., Dumitras, T.: Shallow-deep networks: Understanding and mitigating network overthinking. In: Proceedings of the 36th International Conference on Machine Learning. vol. 97, pp. 3301–3310 (09–15 Jun 2019)
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners (2019)
2019
Earlier work this paper cites.
Beltagy, I., Peters, M.E., Cohan, A.: Longformer: The long-document transformer. arXiv: 2004.05150 (2020)
2020
Cited alongside, same era.
Elbayad, M., Gu, J., Grave, E., Auli, M.: Depth-adaptive transformer. In: International Conference on Learning Representations (2020)
2020
Cited alongside, same era.
Fan, A., Grave, E., Joulin, A.: Reducing transformer depth on demand with structured dropout. In: International Conference on Learning Representations (ICLR) (2020)
2020
Cited alongside, same era.
Hou, L., Huang, Z., Shang, L., Jiang, X., Chen, X., Liu, Q.: Dynabert: Dynamic bert with adaptive width and depth. In: Advances in Neural Information Processing Systems (NeurIPS). vol. 33, pp. 9782–9793 (2020)
2020
Cited alongside, same era.
Kitaev, N., Kaiser, L., Levskaya, A.: Reformer: The efficient transformer. In: International Conference on Learning Representations (2020)
Xin, J., Tang, R., Yu, Y., Lin, J.: BERxiT: Early exiting for BERT with better fine-tuning and extension to regression. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. pp. 91–104 (Apr 2021). https://doi.org/10.18653/v1/2021.eacl-main.8
2021
Later among the works it cites.
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H.W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A.M., Pillai, T.S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., Fiedel, N.: Palm: Scaling language modeling with pathways. arXiv: 2204.02311 (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Lin, B.Y., Zhou, W., Shen, M., Zhou, P., Bhagavatula, C., Choi, Y., Ren, X.: CommonGen: A constrained text generation challenge for generative commonsense reasoning. In: Findings of the Association for Computational Linguistics: EMNLP 2020. pp. 1823–1840 (Nov 2020). https://doi.org/10.18653/v1/2020.findings-emnlp.165
2020
Cited alongside, same era.
Liu, W., Zhou, P., Wang, Z., Zhao, Z., Deng, H., Ju, Q.: FastBERT: a self-distilling BERT with adaptive inference time. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 6035–6044 (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.537
2020
Cited alongside, same era.
Liu, Y., Meng, F., Zhou, J., Chen, Y., Xu, J.: Faster depth-adaptive transformers. In: AAAI Conference on Artificial Intelligence (2020)
2020
Cited alongside, same era.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21
2020
Cited alongside, same era.
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., Smith, N.A.: The right tool for the job: Matching model and instance complexities. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 6640–6651 (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.593
2020
Cited alongside, same era.
Wang, S., Li, B.Z., Khabsa, M., Fang, H., Ma, H.: Linformer: Self-attention with linear complexity. arXiv: 2006.04768 (2020)
2020
Cited alongside, same era.
Xin, J., Tang, R., Lee, J., Yu, Y., Lin, J.: DeeBERT: Dynamic early exiting for accelerating BERT inference. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 2246–2251 (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.204
2020
Cited alongside, same era.
Dao, T., Fu, D.Y., Ermon, S., Rudra, A., Ré, C.: FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In: Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Later among the works it cites.
Hu, E.J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: International Conference on Learning Representations (2022)
2022
Later among the works it cites.
Sajjad, H., Dalvi, F., Durrani, N., Nakov, P.: On the effect of dropping layers of pre-trained transformer models. Computer Speech & Language 77
2022
Later among the works it cites.
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., Metzler, D.: Confident adaptive language modeling. In: Advances in Neural Information Processing Systems (NeurIPS). vol. 35, pp. 17456–17472 (2022)
2022
Later among the works it cites.
Sun, T., Liu, X., Zhu, W., Geng, Z., Wu, L., He, Y., Ni, Y., Xie, G., Huang, X., Qiu, X.: A simple hash-based early exiting approach for language understanding and generation. In: Findings of the Association for Computational Linguistics: ACL 2022. pp. 2409–2421 (May 2022). https://doi.org/10.18653/v1/2022.findings-acl.189
2022
Later among the works it cites.
Xia, M., Zhong, Z., Chen, D.: Structured pruning learns compact and accurate models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 1513–1528 (May 2022). https://doi.org/10.18653/v1/2022.acl-long.107
2022
Later among the works it cites.
Abdin, M., Aneja, J., Bubeck, S., Mendes, C.C.T., Chen, W., Giorno, A.D., Eldan, R., Gopi, S., Gunasekar, S., Javaheripi, M., Kauffmann, P., Lee, Y.T., Li, Y., Nguyen, A., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Santacroce, M., Behl, H.S., Kalai, A.T., Wang, X., Ward, R., Witte, P., Zhang, C., Zhang, Y.: Phi-2: The surprising power of small language models (2023)
2023
Later among the works it cites.
Corro, L.D., Giorno, A.D., Agarwal, S., Yu, B., Awadallah, A., Mukherjee, S.: Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference. arXiv: 2307.02628 (2023)
2023
Later among the works it cites.
Dao, T.: Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv: 2307.08691 (2023)
2023
Later among the works it cites.
Gim, I., Chen, G., seob Lee, S., Sarda, N., Khandelwal, A., Zhong, L.: Prompt cache: Modular attention reuse for low-latency inference. arXiv: 2311.04934 (2023)
2023
Later among the works it cites.
Lan, T., Cai, D., Wang, Y., Huang, H., Mao, X.L.: Copy is all you need. In: The Eleventh International Conference on Learning Representations (2023)
2023
Later among the works it cites.
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K.K., He, X., Hou, H., Lin, J., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K.S.I., Mom, F., Saito, A., Song, G., Tang, X., Wang, B., Wind, J.S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhou, Q., Zhu, J., Zhu, R.J.: Rwkv: Reinventing rnns for the transformer era. arXiv: 2305.13048 (2023)
2023
Later among the works it cites.
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., Wei, F.: Retentive network: A successor to transformer for large language models. arXiv: 2307.08621 (2023)
2023
Later among the works it cites.
Xiao, G., Tian, Y., Chen, B., Han, S., Lewis, M.: Efficient streaming language models with attention sinks. arXiv: 2309.17453 (2023)
2023
Later among the works it cites.