Fetching the paper…
Reading the bibliography…
The burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence.
Pareto, V.: Cours D’économie Politique vol. 1. Librairie Droz, ??? (1964)
1964
Earlier work this paper cites.
Burton, F.W.: Speculative computation, parallelism, and functional programming. IEEE Transactions on Computers 100
1985
Earlier work this paper cites.
LeCun, Y., Denker, J., Solla, S.: Optimal brain damage. Advances in neural information processing systems 2
1989
Earlier work this paper cites.
Vilalta, R., Drissi, Y.: A perspective view and survey of meta-learning. Artificial intelligence review 18
2002
Earlier work this paper cites.
2015
Earlier work this paper cites.
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., Dean, J.: Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In: International Conference on Learning Representations (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Teerapittayanon, S., McDanel, B., Kung, H.-T.: Branchynet: Fast inference via early exiting from deep neural networks. In: 2016 23rd International Conference on Pattern Recognition (ICPR), pp. 2464–2469 (2016). IEEE
2016
Earlier work this paper cites.
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Yepes, A.J., Koehn, P., Logacheva, V., Monz, C., et al
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Shi, W., Cao, J., Zhang, Q., Li, Y., Xu, L.: Edge computing: Vision and challenges. IEEE internet of things journal 3
2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Gunantara, N.: A review of multi-objective optimization: Methods and its applications. Cogent Engineering 5
2018
Earlier work this paper cites.
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al.: Mesh-tensorflow: Deep learning for supercomputers. Advances in neural information processing systems 31
2018
Earlier work this paper cites.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al.: Mixed precision training. The International Conference on Learning Representation (2018)
2018
Earlier work this paper cites.
Johnson, T.B., Guestrin, C.: Training deep models faster with robust, approximate importance sampling. Advances in Neural Information Processing Systems 31
2018
Earlier work this paper cites.
Katharopoulos, A., Fleuret, F.: Not all samples are created equal: Deep learning with importance sampling. In: International Conference on Machine Learning, pp. 2525–2534 (2018). PMLR
2018
Earlier work this paper cites.
Su, D., Zhang, H., Chen, H., Yi, J., Chen, P.-Y., Gao, Y.: Is robustness the cost of accuracy?–a comprehensive study on the robustness of 18 deep image classification models. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 631–648 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Ying, X.: An overview of overfitting and its solutions. In: Journal of Physics: Conference Series, vol. 1168, p. 022022 (2019). IOP Publishing
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Long, X., Ben, Z., Liu, Y.: A survey of related research on compression and acceleration of deep neural networks. In: Journal of Physics: Conference Series, vol. 1213, p. 052003 (2019). IOP Publishing
2019
Earlier work this paper cites.
Kenton, J.D.M.-W.C., Toutanova, L.K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol. 1, p. 2 (2019)
2019
Earlier work this paper cites.
Jang, H., Kim, J., Jo, J.-E., Lee, J., Kim, J.: Mnnfast: A fast and scalable system architecture for memory-augmented neural networks. In: Proceedings of the 46th International Symposium on Computer Architecture, pp. 250–263 (2019)
2019
Earlier work this paper cites.
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q.V., Wu, Y., et al.: Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N.R., Ganger, G.R., Gibbons, P.B., Zaharia, M.: Pipedream: Generalized pipeline parallelism for dnn training. In: Proceedings of the 27th ACM Symposium on Operating Systems Principles, pp. 1–15 (2019)
2019
Earlier work this paper cites.
Song, K., Tan, X., Qin, T., Lu, J., Liu, T.-Y.: Mass: Masked sequence to sequence pre-training for language generation. In: International Conference on Machine Learning, pp. 5926–5936 (2019). PMLR
2019
Earlier work this paper cites.
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., Gelly, S.: Parameter-efficient transfer learning for nlp. In: International Conference on Machine Learning, pp. 2790–2799 (2019). PMLR
2019
Earlier work this paper cites.
Bapna, A., Firat, O.: Simple, scalable adaptation for neural machine translation. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 1538–1548 (2019)
2019
Earlier work this paper cites.
Jawahar, G., Sagot, B., Seddah, D.: What does bert learn about the structure of language? In: ACL 2019-57th Annual Meeting of the Association for Computational Linguistics (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.: Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
Xie, Y., Chen, S., Ni, Q., Wu, H.: Integration of resource allocation and task assignment for optimizing the cost and maximum throughput of business processes. Journal of Intelligent Manufacturing 30
2019
Earlier work this paper cites.
Elsken, T., Metzen, J.H., Hutter, F.: Neural architecture search: A survey. The Journal of Machine Learning Research 20
2019
Earlier work this paper cites.
Floridi, L., Chiriatti, M.: Gpt-3: Its nature, scope, limits, and consequences. Minds and Machines 30
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Rajbhandari, S., Rasley, J., Ruwase, O., He, Y.: Zero: Memory optimizations toward training trillion parameter models. In: SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16 (2020). IEEE
2020
Earlier work this paper cites.
Deng, L., Li, G., Han, S., Shi, L., Xie, Y.: Model compression and hardware acceleration for neural networks: A comprehensive survey. Proceedings of the IEEE 108
2020
Earlier work this paper cites.
Tay, Y., Dehghani, M., Bahri, D., Metzler, D.: Efficient transformers: A survey. arXiv e-prints, 2009 (2020)
2020
Earlier work this paper cites.
Capra, M., Bussolino, B., Marchisio, A., Shafique, M., Masera, G., Martina, M.: An updated survey of efficient hardware architectures for accelerating deep convolutional neural networks. Future Internet 12
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21
2020
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
2020
Earlier work this paper cites.
Kitaev, N., Kaiser, L., Levskaya, A.: Reformer: The efficient transformer. In: International Conference on Learning Representations (2020). https://openreview.net/forum?id=rkgNKkHtvB
2020
Earlier work this paper cites.
Katharopoulos, A., Vyas, A., Pappas, N., Fleuret, F.: Transformers are rnns: Fast autoregressive transformers with linear attention. In: International Conference on Machine Learning, pp. 5156–5165 (2020). PMLR
2020
Earlier work this paper cites.
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., Chen, Z.: Gshard: Scaling giant models with conditional computation and automatic sharding. In: International Conference on Learning Representations (2020)
2020
Earlier work this paper cites.
Jiang, H., He, P., Chen, W., Liu, X., Gao, J., Zhao, T.: SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 2177–2190 (2020)
2020
Earlier work this paper cites.
Xin, J., Tang, R., Lee, J., Yu, Y., Lin, J.: Deebert: Dynamic early exiting for accelerating bert inference. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 2246–2251 (2020)
2020
Earlier work this paper cites.
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., Smith, N.A.: The right tool for the job: Matching model and instance complexities. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6640–6651 (2020)
2020
Earlier work this paper cites.
Zhou, W., Xu, C., Ge, T., McAuley, J., Xu, K., Wei, F.: Bert loses patience: Fast and robust inference with early exit. Advances in Neural Information Processing Systems 33
2020
Earlier work this paper cites.
Iandola, F., Shaw, A., Krishna, R., Keutzer, K.: Squeezebert: What can computer vision teach nlp about efficient neural networks? In: Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, pp. 124–135 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Henderson, P., Hu, J., Romoff, J., Brunskill, E., Jurafsky, D., Pineau, J.: Towards the systematic reporting of the energy and carbon footprints of machine learning. The Journal of Machine Learning Research 21
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Barrault, L., Biesialska, M., Bojar, O., Costa-jussà, M.R., Federmann, C., Graham, Y., Grundkiewicz, R., Haddow, B., Huck, M., Joanis, E., et al
2020
Earlier work this paper cites.
Wang, A., Wolf, T.: Overview of the sustainlp 2020 shared task. In: Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, pp. 174–178 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Menghani, G.: Efficient deep learning: A survey on making deep learning models smaller, faster, and better. ACM Computing Surveys (2021) https://doi.org/10.1145/3578938
2021
Earlier work this paper cites.
Shen, Z., Zhang, M., Zhao, H., Yi, S., Li, H.: Efficient attention: Attention with linear complexities. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3531–3539 (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Wang, X., Xiong, Y., Wei, Y., Wang, M., Li, L.: Lightseq: A high performance inference library for transformers. NAACL-HLT 2021, 113 (2021)
2021
Earlier work this paper cites.
NVIDIA: FasterTransformer: A Faster Transformer Framework. https://github.com/NVIDIA/FasterTransformer (2021)
2021
Earlier work this paper cites.
Baines, M., Bhosale, S., Caggiano, V., Goyal, N., Goyal, S., Ott, M., Lefaudeux, B., Liptchinsky, V., Rabbat, M., Sheiffer, S., et al.: Fairscale: A general purpose modular pytorch library for high performance and large scale training (2021)
2021
Earlier work this paper cites.
Paul, M., Ganguli, S., Dziugaite, G.K.: Deep learning on a data diet: Finding important examples early in training. Advances in Neural Information Processing Systems 34
2021
Earlier work this paper cites.
Xu, R., Luo, F., Zhang, Z., Tan, C., Chang, B., Huang, S., Huang, F.: Raise a child in large language model: Towards effective and generalizable fine-tuning. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9514–9528 (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Hu, E.J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al
2021
Earlier work this paper cites.
Aghajanyan, A., Gupta, S., Zettlemoyer, L.: Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 7319–7328 (2021)
2021
Earlier work this paper cites.
Evans, R.D., Aamodt, T.: Ac-gc: Lossy activation compression with guaranteed convergence. Advances in Neural Information Processing Systems 34
2021
Earlier work this paper cites.
Muhamed, A., Keivanloo, I., Perera, S., Mracek, J., Xu, Y., Cui, Q., Rajagopalan, S., Zeng, B., Chilimbi, T.: Ctr-bert: Cost-effective knowledge distillation for billion-parameter teacher models. In: NeurIPS Efficient Natural Language and Speech Processing Workshop (2021)
2021
Earlier work this paper cites.
Chen, P., Yu, H.-F., Dhillon, I., Hsieh, C.-J.: Drone: Data-aware low-rank compression for large nlp models. Advances in neural information processing systems 34
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Rogers, A., Kovaleva, O., Rumshisky, A.: A primer in bertology: What we know about how bert works. Transactions of the Association for Computational Linguistics 8
2021
Earlier work this paper cites.
Wang, H., Zhang, Z., Han, S.: Spatten: Efficient sparse attention architecture with cascade token and head pruning. In: 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp. 97–110 (2021). IEEE
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Sankar, C., Ravi, S., Kozareva, Z.: Proformer: Towards on-device lsh projection based transformers. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 2823–2828 (2021)
2021
Earlier work this paper cites.
Huang, Z., Hou, L., Shang, L., Jiang, X., Chen, X., Liu, Q.: Ghostbert: Generate more features with cheap operations for bert. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 6512–6523 (2021)
2021
Earlier work this paper cites.
Ma, Z., Ethayarajh, K., Thrush, T., Jain, S., Wu, L., Jia, R., Potts, C., Williams, A., Kiela, D.: Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking. Advances in Neural Information Processing Systems 34
2021
Earlier work this paper cites.
Schmidt, V., Goyal, K., Joshi, A., Feld, B., Conell, L., Laskaris, N., Blank, D., Wilson, J., Friedler, S., Luccioni, S.: Codecarbon: estimate and track carbon emissions from machine learning computing. Cited on, 20 (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A.A., Wilson, A.G.: Does knowledge distillation really work? Advances in Neural Information Processing Systems 34
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Min, S., Boyd-Graber, J., Alberti, C., Chen, D., Choi, E., Collins, M., Guu, K., Hajishirzi, H., Lee, K., Palomaki, J., et al
2021
Earlier work this paper cites.
Ford, B.W., Zong, Z.: Portauthority: Integrating energy efficiency analysis into cross-platform development cycles via dynamic program analysis. Sustainable Computing: Informatics and Systems 30
2021
Earlier work this paper cites.
Shamshirband, S., Joloudari, J.H., Shirkharkolaie, S.K., Mojrian, S., Rahmani, F., Mostafavi, S., Mansor, Z.: Game theory and evolutionary optimization approaches applied to resource allocation problems in computing environments: A survey. Mathematical Biosciences and Engineering (2021)
2021
Earlier work this paper cites.
Murshed, M.S., Murphy, C., Hou, D., Khan, N., Ananthanarayanan, G., Hussain, F.: Machine learning at the network edge: A survey. ACM Computing Surveys (CSUR) 54
2021
Earlier work this paper cites.
Zhang, C., Xie, Y., Bai, H., Yu, B., Li, W., Gao, Y.: A survey on federated learning. Knowledge-Based Systems 216
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Tirumala, K., Markosyan, A., Zettlemoyer, L., Aghajanyan, A.: Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Aminabadi, R.Y., Rajbhandari, S., Awan, A.A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Dhilleswararao, P., Boppu, S., Manikandan, M.S., Cenkeramaddi, L.R.: Efficient hardware architectures for accelerating deep neural networks: Survey. IEEE Access (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Guo, C., Tang, J., Hu, W., Leng, J., Zhang, C., Yang, F., Liu, Y., Guo, M., Zhu, Y.: Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization. In: Proceedings of the 50th Annual International Symposium on Computer Architecture, pp. 1–15 (2023)
2023
Later among the works it cites.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Lefaudeux, B., Massa, F., Liskovich, D., Xiong, W., Caggiano, V., Naren, S., Xu, M., Hu, J., Tintore, M., Zhang, S., Labatut, P., Haziza, D.: xFormers: A modular and hackable Transformer modelling library. https://github.com/facebookresearch/xformers (2022)
2022
Cited alongside, same era.
Dao, T., Fu, D., Ermon, S., Rudra, A., Ré, C.: Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Fedus, W., Zoph, B., Shazeer, N.: Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research 23
2022
Cited alongside, same era.
Du, N., Huang, Y., Dai, A.M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A.W., Firat, O., et al
2022
Cited alongside, same era.
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A.M., Le, Q.V., Laudon, J., et al
2022
Cited alongside, same era.
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X.V., Du, J., Iyer, S., Pasunuru, R., et al
2022
Cited alongside, same era.
Clark, A., De Las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., et al
2022
Cited alongside, same era.
Later among the works it cites.
Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D.: Gptq: Accurate post-training quantization for generative pre-trained transformers. The International Conference on Learning Representations (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Yang, G., Lo, D., Mullins, R., Zhao, Y.: Dynamic stashing quantization for efficient transformer training. Findings of the Association for Computational Linguistics: EMNLP 2023 (2023)
2023
Later among the works it cites.
Yang, Z., Choudhary, S., Kunzmann, S., Zhang, Z.: Quantization-aware and tensor-compressed training of transformers for natural language understanding. Interspeech (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Gong, Z., Liu, J., Wang, Q., Yang, Y., Wang, J., Wu, W., Xian, Y., Zhao, D., Yan, R.: Prequant: A task-agnostic quantization approach for pre-trained language models. Findings of the Association for Computational Linguistics: ACL 2023 (2023)
2023
Later among the works it cites.
Yu, C., Chen, T., Gan, Z.: Boost transformer-based language models with gpu-friendly sparsity and quantization. In: Findings of the Association for Computational Linguistics: ACL 2023, pp. 218–235 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., et al.: Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) (2023)
2023
Later among the works it cites.
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., Hashimoto, T.B.: Stanford alpaca: An instruction-following llama model (2023)
2023
Later among the works it cites.
Ling, C., Zhang, X., Zhao, X., Liu, Y., Cheng, W., Oishi, M., Osaki, T., Matsuda, K., Chen, H., Zhao, L.: Open-ended commonsense reasoning with unrestricted answer candidates. In: Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 8035–8047 (2023)
2023
Later among the works it cites.
Wang, H., Li, R., Jiang, H., Wang, Z., Tang, X., Bi, B., Cheng, M., Yin, B., Wang, Y., Zhao, T., et al
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Tang, S., Wang, Y., Kong, Z., Zhang, T., Li, Y., Ding, C., Wang, Y., Liang, Y., Xu, D.: You need multiple exiting: Dynamic early exiting for accelerating unified vision language model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10781–10791 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Leviathan, Y., Kalman, M., Matias, Y.: Fast inference from transformers via speculative decoding. In: International Conference on Machine Learning, pp. 19274–19286 (2023). PMLR
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Li, S., Liu, H., Bian, Z., Fang, J., Huang, H., Liu, Y., Wang, B., You, Y.: Colossal-ai: A unified deep learning system for large-scale parallel training. In: Proceedings of the 52nd International Conference on Parallel Processing, pp. 766–775 (2023)
2023
Later among the works it cites.
Andonian, A., Anthony, Q., Biderman, S., Black, S., Gali, P., Gao, L., Hallahan, E., Levy-Kramer, J., Leahy, C., Nestler, L., Parker, K., Pieler, M., Phang, J., Purohit, S., Schoelkopf, H., Stander, D., Songz, T., Tigges, C., Thérien, B., Wang, P., Weinbach, S.: GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch (2023). https://doi.org/10.5281/zenodo.5879544 . https://www.github.com/eleutherai/gpt-neox
2023
Later among the works it cites.
Xu, M., Song, C., Tian, Y., Agrawal, N., Granqvist, F., Dalen, R., Zhang, X., Argueta, A., Han, S., Deng, Y., et al
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Wang, Y., Chen, K., Tan, H., Guo, K.: Tabi: An efficient multi-level inference system for large language models. In: Proceedings of the Eighteenth European Conference on Computer Systems, pp. 233–248 (2023)
2023
Later among the works it cites.
Peng, Z., Wang, Z., Deng, D.: Near-duplicate sequence search at scale for large language model memorization evaluation. Proceedings of the ACM on Management of Data 1
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Treviso, M., Lee, J.-U., Ji, T., Aken, B.v., Cao, Q., Ciosici, M.R., Hassid, M., Heafield, K., Hooker, S., Raffel, C., et al
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Jiang, Z., Lin, H., Zhong, Y., Huang, Q., Chen, Y., Zhang, Z., Peng, Y., Li, X., Xie, C., Nong, S., et al
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Wei, W., Ren, X., Tang, J., Wang, Q., Su, L., Cheng, S., Wang, J., Yin, D., Huang, C.: Llmrec: Large language models with graph augmentation for recommendation. In: Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pp. 806–815 (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Lin, X., Wang, W., Li, Y., Yang, S., Feng, F., Wei, Y., Chua, T.-S.: Data-efficient fine-tuning for llm-based recommendation. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 365–374 (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Tang, Z., Zhu, E.: BrainTransformers: SNN-LLM (2024). https://arxiv.org/abs/2410.14687
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Cho, J., Kim, M., Choi, H., Heo, G., Park, J.: Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale. In: 2024 IEEE International Symposium on Workload Characterization (IISWC), pp. 15–29. IEEE, ??? (2024). https://doi.org/10.1109/iiswc63097.2024.00012 . http://dx.doi.org/10.1109/IISWC63097.2024.00012
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Kim, B., Cha, S., Park, S., Lee, J., Lee, S., Kang, S.-h., So, J., Kim, K., Jung, J., Lee, J.-G., Lee, S., Paik, Y., Kim, H., Kim, J.-S., Lee, W.-J., Ro, Y., Cho, Y., Kim, J.H., Song, J., Yu, J., Lee, S., Cho, J., Sohn, K.: The breakthrough memory solutions for improved performance on llm inference. IEEE Micro 44
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.