Fetching the paper…
Reading the bibliography…
Foundation model (FM) powered agent services are regarded as a promising solution to develop intelligent and personalized applications for advancing toward Artificial General Intelligence (AGI).
T. Wang and K. Cho, “Larger-Context Language Modelling with Recurrent Neural Network,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016, pp. 1319–1329
2016
Earlier work this paper cites.
D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A { \{ Low-Latency } \} online prediction serving system,” in 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) , 2017, pp. 613–627
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
W. Na, S. Jang, Y. Lee, L. Park, N.-N. Dao, and S. Cho, “Frequency resource allocation and interference management in mobile edge computing for an Internet of Things system,” IEEE Internet of Things Journal , vol. 6, no. 3, pp. 4910–4920, 2018
2018
Earlier work this paper cites.
H. Jang, J. Kim, J.-E. Jo, J. Lee, and J. Kim, “Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture , 2019, pp. 250–263
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Gorbachev, M. Fedorov, I. Slavutin, A. Tugarev, M. Fatekhov, and Y. Tarkan, “Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,” in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops , 2019, pp. 0–0
2019
Earlier work this paper cites.
H. Shen, L. Chen, Y. Jin, L. Zhao, B. Kong, M. Philipose, A. Krishnamurthy, and R. Sundaram, “Nexus: A GPU cluster engine for accelerating DNN-based video analysis,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 322–337
2019
Earlier work this paper cites.
C. Avasalcai, C. Tsigkanos, and S. Dustdar, “Decentralized resource auctioning for latency-sensitive edge computing,” in 2019 IEEE international conference on edge computing (EDGE) . IEEE, 2019, pp. 72–76
2019
Earlier work this paper cites.
L. Yang, B. Liu, J. Cao, Y. Sahni, and Z. Wang, “Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,” IEEE Transactions on Services Computing , vol. 14, no. 5, pp. 1439–1452, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Zhou, H. Wen, R. Teodorescu, and D. H. Du, “Distributing deep neural networks with containerized partitions at the edge,” in 2nd USENIX Workshop on Hot Topics in Edge Computing (HotEdge 19) , 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC) . IEEE, 2020, pp. 84–89
2020
Earlier work this paper cites.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J.-H. Park, S. Lee, K. Park, J. W. Lee et al. , “Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 328–341
2020
Earlier work this paper cites.
H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte, “Att: A fault-tolerant reram accelerator for attention-based neural networks,” in 2020 IEEE 38th International Conference on Computer Design (ICCD) . IEEE, 2020, pp. 213–221
2020
Earlier work this paper cites.
D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov, “InferLine: latency-aware provisioning and scaling for prediction serving pipelines,” in Proceedings of the 11th ACM Symposium on Cloud Computing , 2020, pp. 477–491
2020
Earlier work this paper cites.
A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y. Vigfusson, and J. Mace, “Serving { \{ DNNs } \} like clockwork: Performance predictability from the bottom up,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) , 2020, pp. 443–462
2020
Earlier work this paper cites.
Z. Tong, X. Deng, F. Ye, S. Basodi, X. Xiao, and Y. Pan, “Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,” Information Sciences , vol. 537, pp. 116–131, 2020
2020
Earlier work this paper cites.
X. Xiong, K. Zheng, L. Lei, and L. Hou, “Resource allocation based on deep reinforcement learning in IoT edge computing,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 6, pp. 1133–1146, 2020
2020
Earlier work this paper cites.
Z. Zhou, S. Yu, W. Chen, and X. Chen, “CE-IoT: Cost-effective cloud-edge resource provisioning for heterogeneous IoT applications,” IEEE Internet of Things Journal , vol. 7, no. 9, pp. 8600–8614, 2020
2020
Earlier work this paper cites.
Z. Chang, L. Liu, X. Guo, and Q. Sheng, “Dynamic resource allocation and computation offloading for IoT fog computing system,” IEEE Transactions on Industrial Informatics , vol. 17, no. 5, pp. 3348–3357, 2020
2020
Earlier work this paper cites.
T. Mohammed, C. Joe-Wong, R. Babbar, and M. Di Francesco, “Distributed inference acceleration with adaptive DNN partitioning and offloading,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications . IEEE, 2020, pp. 854–863
2020
Earlier work this paper cites.
L. Zeng, X. Chen, Z. Zhou, L. Yang, and J. Zhang, “Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,” IEEE/ACM Transactions on Networking , vol. 29, no. 2, pp. 595–608, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Esmaeilzadeh, “Use of AI-based tools for healthcare purposes: a survey study from consumers’ perspectives,” BMC medical informatics and decision making , vol. 20, pp. 1–19, 2020
2020
Earlier work this paper cites.
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in International Conference on Machine Learning . PMLR, 2020, pp. 3690–3699
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. J. Ham, Y. Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2021, pp. 692–705
2021
Earlier work this paper cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2021, pp. 97–110
2021
Earlier work this paper cites.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , 2021, pp. 977–991
2021
Earlier work this paper cites.
A. F. Laguna, A. Kazemi, M. Niemier, and X. S. Hu, “In-memory computing based accelerator for transformer networks for long sequences,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2021, pp. 1839–1844
2021
Earlier work this paper cites.
H. Wang, Z. Qu, S. Guo, N. Wang, R. Li, and W. Zhuang, “LOSP: Overlap synchronization parallel with local compensation for fast distributed training,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 8, pp. 2541–2557, 2021
2021
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Earlier work this paper cites.
F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “ { \{ INFaaS } \} : Automated model-less inference serving,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 397–411
2021
Earlier work this paper cites.
L. Wang, L. Yang, Y. Yu, W. Wang, B. Li, X. Sun, J. He, and L. Zhang, “Morphling: Fast, near-optimal auto-configuration for cloud-native model serving,” in Proceedings of the ACM Symposium on Cloud Computing , 2021, pp. 639–653
2021
Earlier work this paper cites.
B. Wang, A. Ali-Eldin, and P. Shenoy, “Lass: Running latency sensitive serverless computations at the edge,” in Proceedings of the 30th international symposium on high-performance parallel and distributed computing , 2021, pp. 239–251
2021
Earlier work this paper cites.
O. Ascigil, A. G. Tasiopoulos, T. K. Phan, V. Sourlas, I. Psaras, and G. Pavlou, “Resource provisioning and allocation in function-as-a-service edge-clouds,” IEEE Transactions on Services Computing , vol. 15, no. 4, pp. 2410–2424, 2021
2021
Earlier work this paper cites.
J. Li, W. Liang, Y. Li, Z. Xu, X. Jia, and S. Guo, “Throughput maximization of delay-aware DNN inference in edge computing by exploring DNN model partitioning and inference parallelism,” IEEE Transactions on Mobile Computing , vol. 22, no. 5, pp. 3017–3030, 2021
2021
Earlier work this paper cites.
W. Zeng, X. Ren, T. Su, H. Wang, Y. Liao, Z. Wang, X. Jiang, Z. Yang, K. Wang, X. Zhang et al. , “Pangu- α \alpha : Large-scale autoregressive pretrained chinese language models with auto-parallel computation,” arXiv e-prints , pp. arXiv–2104, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham, “Jurassic-1: Technical details and evaluation,” White Paper. AI21 Labs , vol. 1, p. 9, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Schuster, A. Fisch, T. Jaakkola, and R. Barzilay, “Consistent Accelerated Inference via Confident Adaptive Transformers,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 4962–4979
2021
Earlier work this paper cites.
G. Kim and K. Cho, “Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 6501–6511
2021
Earlier work this paper cites.
Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” Advances in neural information processing systems , vol. 34, pp. 13 937–13 949, 2021
2021
Earlier work this paper cites.
Y. Liang, G. Chongjian, Z. Tong, Y. Song, J. Wang, and P. Xie, “EViT: Expediting Vision Transformers via Token Reorganizations,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 389–402
2021
Earlier work this paper cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al. , “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
OpenAI, “Openai codex,” 2021. [Online]. Available: https://openai.com/index/openai-codex/
2021
Earlier work this paper cites.
H. Djigal, J. Xu, L. Liu, and Y. Zhang, “Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,” IEEE Communications Surveys & Tutorials , vol. 24, no. 4, pp. 2449–2494, 2022
2022
Earlier work this paper cites.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2022, pp. 616–630
2022
Earlier work this paper cites.
Z. Zhou, J. Liu, Z. Gu, and G. Sun, “Energon: Toward efficient acceleration of transformers using dynamic sparse attention,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 42, no. 1, pp. 136–149, 2022
2022
Earlier work this paper cites.
J. Choi, H. Li, B. Kim, S. Hwang, and J. H. Ahn, “Accelerating transformer networks through recomposing softmax layers,” in 2022 IEEE International Symposium on Workload Characterization (IISWC) . IEEE, 2022, pp. 92–103
2022
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Earlier work this paper cites.
X. Han, G. Zeng, W. Zhao, Z. Liu, Z. Zhang, J. Zhou, J. Zhang, J. Chao, and M. Sun, “Bminf: An efficient toolkit for big model inference and tuning,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , 2022, pp. 224–230
2022
Earlier work this paper cites.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in International conference on machine learning . PMLR, 2022, pp. 18 332–18 346
2022
Earlier work this paper cites.
X. Mu, Y. Liu, L. Guo, and N. Al-Dhahir, “Heterogeneous semantic and bit communications: A semi-noma scheme,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 1, pp. 155–169, 2022
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Earlier work this paper cites.
Gunasekaran, Jashwant Raj and Mishra, Cyan Subhra and Thinakaran, Prashanth and Sharma, Bikash and Kandemir, Mahmut Taylan and Das, Chita R, “Cocktail: A multidimensional optimization for model serving in cloud,” in 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) , 2022, pp. 1041–1057
2022
Earlier work this paper cites.
X. Li, P. Kang, J. Molone, W. Wang, and P. Lama, “KneeScale: Efficient resource scaling for serverless computing at the edge,” in 2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid) . IEEE, 2022, pp. 180–189
2022
Earlier work this paper cites.
S. Hu, W. Shi, and G. Li, “CEC: A containerized edge computing framework for dynamic resource provisioning,” IEEE Transactions on Mobile Computing , 2022
2022
Earlier work this paper cites.
Y. Wu, M. Lentz, D. Zhuo, and Y. Lu, “Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures,” Proceedings of the VLDB Endowment , vol. 16, no. 3, pp. 406–419, 2022
2022
Earlier work this paper cites.
Y. Hu, C. Imes, X. Zhao, S. Kundu, P. A. Beerel, S. P. Crago, and J. P. Walters, “PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices,” in 2022 25th Euromicro Conference on Digital System Design (DSD) , 2022, pp. 298–307
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” 2022
2022
Earlier work this paper cites.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei, “Scaling instruction-finetuned language models,” 2022
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton, “Pix2seq: A language modeling framework for object detection,” 2022
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
Earlier work this paper cites.
K. Zhao, Z. Zhou, X. Chen, R. Zhou, X. Zhang, S. Yu, and D. Wu, “Edgeadaptor: Online configuration adaption, model selection and resource provisioning for edge dnn inference serving at scale,” IEEE Transactions on Mobile Computing , 2022
2022
Earlier work this paper cites.
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer, “Learned token pruning for transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 784–794
2022
Earlier work this paper cites.
A. Modarressi, H. Mohebbi, and M. T. Pilehvar, “AdapLeR: Speeding up Inference by Adaptive Length Reduction,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 1–15
2022
Earlier work this paper cites.
Y. Tang, K. Han, Y. Wang, C. Xu, J. Guo, C. Xu, and D. Tao, “Patch slimming for efficient vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 165–12 174
2022
Earlier work this paper cites.
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token Merging: Your ViT But Faster,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
J. Wang, X. Yang, H. Li, L. Liu, Z. Wu, and Y.-G. Jiang, “Efficient video transformers with spatial-temporal token selection,” in European Conference on Computer Vision . Springer, 2022, pp. 69–86
2022
Earlier work this paper cites.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient Transformers: A Survey,” ACM Comput. Surv. , vol. 55, no. 6, dec 2022
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Earlier work this paper cites.
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huang, and T. Gui, “The Rise and Potential of Large Language Model Based Agents: A Survey,” 2023
2023
Earlier work this paper cites.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, H. Jin, T. Chen, and Z. Jia, “Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,” 2023
2023
Earlier work this paper cites.
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang, R. Ren, Y. Li, X. Tang, Z. Liu, P. Liu, J.-Y. Nie, and J.-R. Wen, “A Survey of Large Language Models,” 2023
2023
Earlier work this paper cites.
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang, “A Survey on Model Compression for Large Language Models,” 2023
2023
Earlier work this paper cites.
Y. Bai, H. Zhou, K. Zhao, J. Chen, J. Yu, and K. Wang, “Transformer-opu: An fpga-based overlay processor for transformer networks,” in 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . IEEE, 2023, pp. 221–221
2023
Earlier work this paper cites.
B. Reidy, M. Mohammadi, M. Elbtity, H. Smith, and Z. Ramtin, “Work in progress: Real-time transformer inference on edge ai accelerators,” in 2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS) , 2023, pp. 341–344
2023
Earlier work this paper cites.
B. He and T. Hofmann, “Simplifying transformer blocks,” arXiv preprint arXiv:2311.01906 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Re, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” 2023
2023
Earlier work this paper cites.
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Re et al. , “Deja vu: Contextual sparsity for efficient llms at inference time,” in International Conference on Machine Learning . PMLR, 2023, pp. 22 137–22 176
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai, “Gqa: Training generalized multi-query transformer models from multi-head checkpoints,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y. Y. Wong, A. Zhu, L. Yang, X. Shi, C. Shi, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia, “Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Hong, X. Qiu, J. Lin, W. Chen, Y. Yu, H. Wang, S. Guo, and W. Gao, “Intelligence-endogenous management platform for computing and network convergence,” IEEE Network , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Q. Cao, B. Paranjape, and H. Hajishirzi, “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 12 890–12 903
2023
Later among the works it cites.
S. Ren, S. Chen, S. Li, X. Sun, and L. Hou, “TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 932–947
2023
Later among the works it cites.
S. Ding, P. Zhao, X. Zhang, R. Qian, H. Xiong, and Q. Tian, “Prune spatio-temporal tokens by semantic-aware temporal accumulation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 16 945–16 956
2023
Later among the works it cites.
J. B. Haurum, S. Escalera, G. W. Taylor, and T. B. Moeslund, “Which tokens to use? investigating token reduction in vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 773–783
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Gerganov, “ggerganov/llama.cpp: Port of facebook’s llama model in c/c++.” https://github.com/ggerganov/llama.cpp , 2023
2023
Cited alongside, same era.
M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/mlc-ai/mlc-llm
2023
Cited alongside, same era.
mnn llm, “mnn-llm: llm deploy project based mnn.” https://github.com/wangzhaode/mnn-llm , 2023
2023
Cited alongside, same era.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
Cited alongside, same era.
mllm, “mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices.” https://github.com/UbiquitousLearning/mllm , 2023
2023
Cited alongside, same era.
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y. Liu, B. Wang, and Y. You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing , 2023, pp. 766–775
2023
Cited alongside, same era.
tensorrtllm, “A tensorrt toolbox for optimized large language model inference,” https://github.com/NVIDIA/TensorRT-LLM , 2023
2023
Cited alongside, same era.
B. Li, S. Samsi, V. Gadepally, and D. Tiwari, “Kairos: Building cost-efficient machine learning inference systems with heterogeneous cloud resources,” in Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing , 2023, pp. 3–16
2023
Cited alongside, same era.
2023
Later among the works it cites.
F. Xue, V. Likhosherstov, A. Arnab, N. Houlsby, M. Dehghani, and Y. You, “Adaptive computation with elastic input sequence,” in International Conference on Machine Learning . PMLR, 2023, pp. 38 971–38 988
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y. Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, and B. Ichter, “Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents,” in Conference on Neural Information Processing Systems , 2023
2023
Later among the works it cites.
A. Omidvar and A. An, “Empowering Conversational Agents using Semantic In-Context Learning,” in Annual Meeting of the Association for Computational Linguistics , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Lee, V. Hartmann, J. Park, D. Papailiopoulos, and K. Lee, “Prompted LLMs as Chatbot Modules for Long Open-domain Conversation,” in Annual Meeting of the Association for Computational Linguistics , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Hatalis, D. Christou, J. Myers, S. Jones, K. Lambert, A. Amos-Binks, Z. Dannenhauer, and D. Dannenhauer, “Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents,” in Proceedings of the AAAI Symposium Series , vol. 2, no. 1, 2023, pp. 277–280
2023
Later among the works it cites.
X. Li and X. Qiu, “Mot: Memory-of-thought enables chatgpt to self-improve,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 6354–6374
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhuang, Y. Yu, K. Wang, H. Sun, and C. Zhang, “ToolQA: A Dataset for LLM Question Answering with External Tools,” in Advances in Neural Information Processing Systems , A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 50 117–50 143. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/file/9cb2a7495900f8b602cb10159246a016-Paper-Datasets_and_Benchmarks.pdf
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Hao, T. Liu, Z. Wang, and Z. Hu, “ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings,” in Advances in Neural Information Processing Systems , A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 45 870–45 894. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/file/8fd1a81c882cd45f64958da6284f4a3f-Paper-Conference.pdf
2023
Later among the works it cites.
B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro, “ART: Automatic multi-step reasoning and tool-use for large language models,” 2023
2023
Later among the works it cites.
Y. Wen and S. Chaudhuri, “Batched Low-Rank Adaptation of Foundation Models,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
(2023) Github copilot: Your ai pair programmer. Github. [Online]. Available: https://github.com/features/copilot
2023
Later among the works it cites.
Midjourney, “Midjourney,” 2023. [Online]. Available: https://www.midjourney.com/
2023
Later among the works it cites.
Runway, “Gen-2 by runway,” 2023. [Online]. Available: https://research.runwayml.com/gen2
2023
Later among the works it cites.
Toran Bruce Richards, Pi, Blake Werlinger, Douglas Schonholtz, Hunter Araujo, Dion, David Wurtz, Fergus, Andrew Minnella, Ian, Robin Sallay, “Auto-GPT,” https://news.agpt.co/ , 2023
2023
Later among the works it cites.
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” in Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , 2023, pp. 1–22
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Nerdynav, “107 Up-to-Date ChatGPT Statistics & User Numbers,” 2024, accessed: 2024-04-24. [Online]. Available: https://nerdynav.com/chatgpt-statistics/
2024
Closest in time.
C. Kachris, “A Survey on Hardware Accelerators for Large Language Models,” 2024
2024
Closest in time.
S. Tang, Y. Yu, H. Wang, G. Wang, W. Chen, Z. Xu, S. Guo, and W. Gao, “A Survey on Scheduling Techniques in Computing and Network Convergence,” IEEE Communications Surveys & Tutorials , vol. 26, no. 1, pp. 160–195, 2024
2024
Closest in time.
M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, V. C. M. Leung, and H. V. Poor, “Unleashing the Power of Edge-Cloud Generative AI in Mobile Networks: A Survey of AIGC Services,” IEEE Communications Surveys & Tutorials , pp. 1–1, 2024
2024
Closest in time.
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, “A Comprehensive Overview of Large Language Models,” 2024
2024
Closest in time.
W. Wang, W. Chen, Y. Luo, Y. Long, Z. Lin, L. Zhang, B. Lin, D. Cai, and X. He, “Model Compression and Efficient Inference for Large Language Models: A Survey,” 2024
2024
Closest in time.
2024
Closest in time.
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, pp. 1–26, 2024
2024
Closest in time.
2024
Closest in time.
H. Wang, Y. Jia, M. Zhang, Q. Hu, H. Ren, P. Sun, Y. Wen, and T. Zhang, “FedDSE: Distribution-aware Sub-model Extraction for Federated Learning over Resource-constrained Devices,” in Proceedings of the ACM on Web Conference 2024 , 2024, pp. 2902–2913
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Qin, J. Ying, D. Yang, H. Wang, and X. Tao, “Computing networks enabled semantic communications,” IEEE Network , 2024
2024
Closest in time.
2024
Closest in time.
Harrison Chase, “Langchain,” https://github.com/langchain-ai/langchain , 2024, accessed: 2024-04-07
2024
Closest in time.
C. Lin, Z. Han, C. Zhang, Y. Yang, F. Yang, C. Chen, and L. Qiu, “Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Association, Jul. 2024
2024
Closest in time.
L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng, “SGLang: Efficient Execution of Structured Language Model Programs,” 2024
2024
Closest in time.
H. Wang, H. Xu, Y. Li, Y. Xu, R. Li, and T. Zhang, “FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y. Belkada, P. Samygin, and C. A. Raffel, “Distributed Inference and Fine-tuning of Large Language Models Over The Internet,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
H. Li, X. Li, Q. Fan, Q. He, X. Wang, and V. C. Leung, “Distributed DNN Inference with Fine-grained Model Partitioning in Mobile Edge Computing Networks,” IEEE Transactions on Mobile Computing , 2024
2024
Closest in time.
Z. Liu, M. Tian, M. Dong, X. Wang, C. Qiu, and C. Zhang, “MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration,” IEEE Transactions on Mobile Computing , no. 01, pp. 1–14, 2024
2024
Closest in time.
NVIDIA Corporation, “NVIDIA Triton Inference Server,” https://developer.nvidia.com/nvidia-triton-inference-server , 2024, accessed: 2024-04-17
2024
Closest in time.
Google LLC, “TensorFlow Serving,” https://www.tensorflow.org/tfx/guide/serving , 2024, accessed: 2024-04-17
2024
Closest in time.
H. Wang, P. Zheng, X. Han, W. Xu, R. Li, and T. Zhang, “FedNLR: Federated Learning with Neuron-wise Learning Rates,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 3069–3080
2024
Closest in time.
BaiChuan-Inc, “Baichuan-7b: About: A large-scale 7b pretraining language model developed by baichuan-inc,” https://github.com/baichuan-inc/Baichuan-7B/tree/main , 2024, accessed: 2024-04-07
2024
Closest in time.
Baichuan Intelligent Technology, “A 13b large language model developed by baichuan intelligent technology,” https://github.com/baichuan-inc/Baichuan-13B/tree/main , 2024, accessed: 2024-04-07
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
S. Anagnostidis, D. Pavllo, L. Biggio, L. Noci, A. Lucchi, and T. Hofmann, “Dynamic context pruning for efficient and interpretable autoregressive transformers,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
E. Kurtić, E. Frantar, and D. Alistarh, “ZipLM: Inference-Aware Structured Pruning of Language Models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
J. Chee, Y. Cai, V. Kuleshov, and C. M. De Sa, “Quip: 2-bit quantization of large language models with guarantees,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
C. Zhang, Y. Yang, Q. Wang, J. Liu, J. Wang, W. Wu, and D. Song, “Minimal Distillation Schedule for Extreme Language Model Compression,” in Findings of the Association for Computational Linguistics: EACL 2024 , 2024, pp. 1378–1394
2024
Closest in time.
V. Kontonis, F. Iliopoulos, K. Trinh, C. Baykal, G. Menghani, and E. Vee, “Slam: Student-label mixing for distillation with unlabeled examples,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Mu, X. Li, and N. Goodman, “Learning to compress prompts with gist tokens,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Chen, W. Xu, Z. Hong, S. Guo, H. Wang, J. Zhang, and D. Zeng, “OTAS: An Elastic Transformer Serving System via Token Adaptation,” 2024
2024
Closest in time.
W. Chen, Y. Su, J. Zuo, C. Yang, C. Yuan, C.-M. Chan, yang Yu, Y. Lu, Y.-H. Hung, C. Qian, Y. Qin, X. Cong, R. Xie, Z. Liu, M. Sun, and J. Zhou, “AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=EHg5GDnyq1
2024
Closest in time.
Z. Wang, S. Cai, G. Chen, A. Liu, X. S. Ma, and Y. Liang, “Describe, explain, plan and select: interactive planning with LLMs enables open-world multi-task agents,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Zhao, W. S. Lee, and D. Hsu, “Large language models as commonsense knowledge for large-scale task planning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
E. Brooks, L. Walls, R. L. Lewis, and S. Singh, “Large language models can implement policy iteration,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
T. Dinh, J. Zhao, S. Tan, R. Negrinho, L. Lausen, S. Zha, and G. Karypis, “Large language models of code fail at completing code with potential bugs,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Leveraging pre-trained large language models to construct and utilize world models for model-based task planning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
D. Zhang, L. Chen, S. Zhang, H. Xu, Z. Zhao, and K. Yu, “Large Language Models Are Semi-Parametric Reinforcement Learning Agents,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei, “Augmenting language models with long-term memory,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. Yang, L. Song, Y. Li, S. Zhao, Y. Ge, X. Li, and Y. Shan, “Gpt4tools: Teaching large language model to use tools via self-instruction,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Zheng, X. Ren, F. Xue, Y. Luo, X. Jiang, and Y. You, “Response length perception and sequence scheduling: An llm-empowered llm inference pipeline,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
“Openai chatgpt,” https://chat.openai.com , accessed: 2024-05-12
2024
Closest in time.
Baidu, “Wenxin yiyan,” 2024. [Online]. Available: https://yiyan.baidu.com/
2024
Closest in time.
OpenAI, “Sora,” 2024. [Online]. Available: https://openai.com/index/sora/
2024
Closest in time.
S. Wei, T. Ye, S. Zhang, Y. Tang, and J. Liang, “Joint token pruning and squeezing towards more aggressive compression of vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2092–2101
2092
Closest in time.