Fetching the paper…
Reading the bibliography…
On-device large language models (LLMs), referring to running LLMs on edge devices, have raised considerable interest since they are more cost-effective, latency-efficient, and privacy-preserving compared with the cloud paradigm.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 33, Dec. 2020, pp. 1877–1901
1901
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” C Users Journal , vol. 12, pp. 23–38, Feb. 1994
1994
Earlier work this paper cites.
M. Satyanarayanan, P. Bahl, R. Caceres, and N. Davies, “The case for VM-based cloudlets in mobile computing,” IEEE Pervasive Comput. , vol. 8, no. 4, pp. 14–23, Oct.-Dec. 2009
2009
Earlier work this paper cites.
L. Skorin-Kapov and M. Matijasevic, “Analysis of QoS requirements for e-health services and mapping to evolved packet system QoS classes,” International journal of telemedicine and applications , vol. 2010, no. 1, p. 628086, Oct. 2010
2010
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP) , Kyoto, Japan, Mar. 2012, pp. 5149–5152
2012
Earlier work this paper cites.
A. Graves, Long Short-Term Memory . Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 37–45
2012
Earlier work this paper cites.
K. Poularakis and L. Tassiulas, “Exploiting user mobility for wireless content delivery,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT) , Istanbul, Turkey, Jul. 2013, pp. 1017–1021
2013
Earlier work this paper cites.
G. Xylomenos, C. N. Ververidis, V. A. Siris, N. Fotiou, C. Tsilopoulos, X. Vasilakos, K. V. Katsaros, and G. C. Polyzos, “A survey of information-centric networking research,” IEEE Commun. Surveys Tuts. , vol. 16, no. 2, pp. 1024–1049, 2nd Quart. 2014
2014
Earlier work this paper cites.
J. Gu, W. Wang, A. Huang, H. Shan, and Z. Zhang, “Distributed cache replacement for caching-enable base stations in cellular networks,” in Proc. IEEE Int. Conf. Commun. (ICC) , Sydney, Australia, Jun. 2014, pp. 2648–2653
2014
Earlier work this paper cites.
GDPR, “Art. 9 GDPR: Processing of special categories of personal data,” 2015. [Online]. Available: https://gdpr-info.eu/art-9-gdpr/
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , Boston, MA, USA, Jun. 2015, pp. 3156–3164
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, May 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proc. Annu. Meet. Assoc. Comput. Linguistics , Aug. 2016, pp. 1715–1725
2016
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . Cambridge, MA, USA: MIT press, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , Long Beach, CA, USA, Dec. 2017, pp. 5998–6008
2017
Earlier work this paper cites.
G. Ananthanarayanan, P. Bahl, P. Bodík, K. Chintalapudi, M. Philipose, L. Ravindranath, and S. Sinha, “Real-time video analytics: The killer app for edge computing,” Computer , vol. 50, no. 10, pp. 58–67, Oct. 2017
2017
Earlier work this paper cites.
J. Cao, L. Xu, R. Abdallah, and W. Shi, “EdgeOS_H: A home operating system for internet of everything,” in Proc. IEEE Int. Conf. Distrib. Comput. Syst. (ICDCS) , Atlanta, GA, USA, Jun. 2017, pp. 1756–1764
2017
Earlier work this paper cites.
X. Dong, S. Chen, and S. Pan, “Learning to prune deep neural networks via layer-wise optimal brain surgeon,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , Long Beach, CA, USA, Dec. 2017
2017
Earlier work this paper cites.
J. Novikova, O. Dušek, and V. Rieser, “The E2E dataset: New challenges for end-to-end generation,” in Proc. Annu. SIGdial Meeting Disc. Dialogue , Saarbrücken, Germany, Aug. 2017, pp. 201–206
2017
Earlier work this paper cites.
M. Ott, S. Edunov, D. Grangier, and M. Auli, “Scaling neural machine translation,” in Proc. ACL Conf. Mach. Transl. (WMT) , Brussels, Belgium, Oct. 2018, pp. 1–9
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018. [Online]. Available: https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
2018
Earlier work this paper cites.
L. Li, K. Ota, and M. Dong, “Deep learning for smart industry: Efficient manufacture inspection system with fog computing,” IEEE Trans. Ind. Informat. , vol. 14, no. 10, pp. 4665–4673, Oct. 2018
2018
Earlier work this paper cites.
A. Seetharam, “On caching and routing in information-centric networks,” IEEE Commun. Mag. , vol. 56, no. 3, pp. 204–209, Mar. 2018
2018
Earlier work this paper cites.
O. Gupta and R. Raskar, “Distributed learning of deep neural network over multiple agents,” J. Netw. Comput. Appl. , vol. 116, pp. 1–8, Aug. 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguist.: Human Lang. Technol., Vol. 1 , Minneapolis, MN, USA, Jun. 2019, pp. 4171–4186
2019
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , Long Beach, CA, USA, Jun. 2019, pp. 10 502–10 511
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “XLNet: Generalized autoregressive pretraining for language understanding,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , vol. 32, Red Hook, NY, USA, Dec. 2019, pp. 5753–5763
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, pp. 1–24, 2019
2019
Earlier work this paper cites.
X. Zhang, Y. Wang, S. Lu, L. Liu, W. Shi et al. , “OpenEI: An open framework for edge intelligence,” in Proc. IEEE Int. Conf. Distrib. Comput. Syst. (ICDCS) , Dallas, TX, USA, Jul. 2019, pp. 1840–1851
2019
Earlier work this paper cites.
Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proc. IEEE , vol. 107, no. 8, pp. 1738–1762, Aug. 2019
2019
Earlier work this paper cites.
ITU, “Architectural framework for machine learning in future networks including IMT-2020,” International Telecommunication Union (ITU), ITU-T Recommendation ITU-T Y.3172, Jun. 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in Proc. Int. Conf. Mach. Learn. (ICML) , Long Beach, CA, USA, Jun. 2019, pp. 2790–2799
2019
Earlier work this paper cites.
N. H. Tran, W. Bao, A. Zomaya, M. N. H. Nguyen, and C. S. Hong, “Federated learning over wireless networks: Optimization model design and analysis,” in Proc. IEEE Int. Conf. Comput. Commun. (INFOCOM) , Paris, France, Apr. 2019, pp. 1387–1395
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “GPipe: Efficient training of giant neural networks using pipeline parallelism,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , Vancouver, BC, Canada, Dec. 2019, pp. 103–112
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Hu, W. Bao, D. Wang, and F. Liu, “Dynamic adaptive DNN surgery for inference acceleration on the edge,” in Proc. IEEE Int. Conf. Comput. Commun. (INFOCOM) , Paris, France, Apr. 2019, pp. 1423–1431
2019
Earlier work this paper cites.
H. Ding, Y. Guo, X. Li, and Y. Fang, “Beef up the edge: Spectrum-aware placement of edge computing services for the internet of things,” IEEE Trans. Mobile Comput. , vol. 18, no. 12, pp. 2783–2795, Dec. 2019
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. ECVA Eur. Conf. Comput. Vis. (ECCV) , Aug. 2020, pp. 213–229
2020
Earlier work this paper cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A lite BERT for self-supervised learning of language representations,” in Proc. Int. Conf. Learn. Represent. , Apr. 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 140, pp. 1–67, Jun. 2020
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , Jul. 2020, pp. 7871–7880
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , vol. 33, Dec. 2020, pp. 9459–9474
2020
Earlier work this paper cites.
C. Wu, K. He, J. Chen, Z. Zhao, and R. Du, “Liveness is not enough: Enhancing fingerprint authentication with behavioral biometrics to defeat puppet attacks,” in Proc. USENIX Secur. Symp. , Aug. 2020, pp. 2219–2236
2020
Earlier work this paper cites.
J. Wang, C. Jiang, H. Zhang, Y. Ren, K.-C. Chen, and L. Hanzo, “Thirty years of machine learning: The road to pareto-optimal wireless networks,” IEEE Commun. Surveys Tuts. , vol. 22, no. 3, pp. 1472–1514, 3rd Quart. 2020
2020
Earlier work this paper cites.
C. Wu, K. He, J. Chen, R. Du, and Y. Xiang, “CaIAuth: Context-aware implicit authentication when the screen is awake,” IEEE Internet Things J. , vol. 7, no. 12, pp. 11 420–11 430, Dec. 2020
2020
Earlier work this paper cites.
S. Deng, H. Zhao, W. Fang, J. Yin, S. Dustdar, and A. Y. Zomaya, “Edge intelligence: The confluence of edge computing and artificial intelligence,” IEEE Internet Things J. , vol. 7, no. 8, pp. 7457–7469, Aug. 2020
2020
Earlier work this paper cites.
L. Soldaini and A. Moschitti, “The cascade transformer: An application for efficient answer sentence selection,” in Proc. Annu. Meet. Assoc. Comput. Linguist. (ACL) , Jul. 2020, pp. 5697–5708
2020
Earlier work this paper cites.
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “DeeBERT: Dynamic early exiting for accelerating BERT inference,” in Proc. Annu. Meet. Assoc. Comput. Linguist. (ACL) , Jul. 2020, pp. 2246–2251
2020
Earlier work this paper cites.
C. Zhong, M. C. Gursoy, and S. Velipasalar, “Deep reinforcement learning-based edge caching in wireless networks,” IEEE Trans. on Cogn. Commun. Netw. , vol. 6, no. 1, pp. 48–61, Mar. 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “ZeRO: Memory optimizations toward training trillion parameter models,” in Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal. , Atlanta, GA, USA, Nov. 2020, pp. 1–16
2020
Earlier work this paper cites.
Y. Liu, Z. Zeng, W. Tang, and F. Chen, “Data-importance aware radio resource allocation: Wireless communication helps machine learning,” IEEE Commun. Lett. , vol. 24, no. 9, pp. 1981–1985, May 2020
2020
Earlier work this paper cites.
L. Liu, J. Zhang, S. Song, and K. B. Letaief, “Client-edge-cloud hierarchical federated learning,” in Proc. IEEE Int. Conf. Commun. (ICC) , Dublin, Ireland, Jun. 2020, pp. 1–6
2020
Earlier work this paper cites.
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma, “PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination,” in Proc. Int. Conf. Mach. Learn. (ICML) , Jul. 2020, pp. 3690–3699
2020
Earlier work this paper cites.
3GPP, “3rd generation partnership project; Technical specification group services and system aspects; Study on traffic characteristics and performance requirements for AI/ML model transfer in 5GS; (Release 18),” 3rd Generation Partnership Project (3GPP), Technical Specification (TS) 22.874, Dec. 2021, version 18.2.0
2021
Earlier work this paper cites.
Nvidia, “Creating voice-based virtual assistants using NVIDIA riva and rasa,” 2021. [Online]. Available: https://developer.nvidia.com/blog/creating-voice-based-virtual-assistants-using-nvidia-riva-and-rasa/
2021
Earlier work this paper cites.
G. Soyalp, A. Alar, K. Ozkanli, and B. Yildiz, “Improving text classification with transformer,” in Proc. IEEE Int. Conf. Comput. Sci. Eng. (UBMK) , Sep. 2021, pp. 707–712
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. (ICLR) , May 2021, pp. 1–21
2021
Earlier work this paper cites.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Trans. Intell. Syst. Technol. , vol. 12, no. 5, pp. 1–32, Oct. 2021
2021
Earlier work this paper cites.
R. Dale, “GPT-3: What’s it good for?” Nat. Lang. Eng. , vol. 27, no. 1, pp. 113–118, Jan. 2021
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Proc. Int. Conf. Mach. Learn. , Jul. 2021, pp. 10 347–10 357
2021
Earlier work this paper cites.
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mT5: A massively multilingual pre-trained text-to-text transformer,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol. , Jun. 2021, pp. 483–498
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , vol. 29, p. 3451–3460, Oct. 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
C. Wu, K. He, J. Chen, Z. Zhao, and R. Du, “Toward robust detection of puppet attacks via characterizing fingertip-touch behaviors,” IEEE Trans. Dependable Secure. Comput. , vol. 19, no. 6, pp. 4002–4018, Nov.-Dec. 2021
2021
Earlier work this paper cites.
Z. Lin, S. Bi, and Y.-J. A. Zhang, “Optimizing AI service placement and resource allocation in mobile edge intelligence systems,” IEEE Trans. Wireless Commun. , vol. 20, no. 11, pp. 7257–7271, Nov. 2021
2021
Earlier work this paper cites.
D. Xu, T. Li, Y. Li, X. Su, S. Tarkoma, T. Jiang, J. Crowcroft, and P. Hui, “Edge intelligence: Empowering intelligence to the edge of network,” Proc. IEEE , vol. 109, no. 11, pp. 1778–1837, Nov. 2021
2021
Earlier work this paper cites.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP) , Online and Punta Cana, Dominican Republic, Nov. 2021, pp. 3045–3059
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proc. 59th Annu. Meet. Assoc. Comput. Linguist. and 11th Int. Joint Conf. Nat. Lang. Process. , Aug. 2021, pp. 4582–4597
2021
Earlier work this paper cites.
D. Guo, A. M. Rush, and Y. Kim, “Parameter-efficient transfer learning with diff pruning,” in Proc. 59th Annu. Meet. Assoc. Comput. Linguist. and 11th Int. Joint Conf. Nat. Lang. Process , Aug. 2021, pp. 4884–4896
2021
Earlier work this paper cites.
X. Li, Y. Shao, T. Sun, H. Yan, X. Qiu, and X. Huang, “Accelerating BERT inference for sequence labeling via early-exit,” in Proc. 59th Annu. Meet. Assoc. Comput. Linguist. and 11th Int. Joint Conf. Nat. Lang. Process. , Aug. 2021, pp. 189–199
2021
Earlier work this paper cites.
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen, “GShard: Scaling giant models with conditional computation and automatic sharding,” in Proc. Int. Conf. Learn. Represent. (ICLR) , may 2021, pp. 1–23
2021
Earlier work this paper cites.
Huawei, 6G: The Next Horizon: From Connected People and Things to Connected Intelligence . Cambridge, U.K.: Cambridge Univ. Press, 2021
2021
Earlier work this paper cites.
N. Hudson, H. Khamfroush, and D. E. Lucani, “QoS-aware placement of deep learning services on the edge with multiple service implementations,” in Proc. Int. Conf. on Comput. Commun. and Netw. (ICCCN) , Athens, Greece, Jul. 2021, pp. 1–8
2021
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on GPU clusters using megatron-LM,” in Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal. , New York, NY, USA, Nov. 2021, pp. 1–15
2021
Earlier work this paper cites.
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He, “ZeRO-Offload: Democratizing billion-scale model training,” in Proc. USENIX Annu. Tech. Conf. (USENIX ATC) , Jul. 2021, pp. 551–564
2021
Earlier work this paper cites.
D. Liu, G. Zhu, J. Zhang, and K. Huang, “Data-importance aware user scheduling for communication-efficient edge machine learning,” IEEE Trans. on Cogn. Commun. Netw. , vol. 7, no. 1, pp. 265–278, Mar. 2021
2021
Earlier work this paper cites.
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He, “ZeRO-Infinity: Breaking the GPU memory wall for extreme scale deep learning,” in Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal. , St. Louis, MO, USA, Nov. 2021, pp. 1–14
2021
Earlier work this paper cites.
S. Wang, X. Zhang, H. Uchiyama, and H. Matsuda, “HiveMind: Towards cellular native machine learning model splitting,” IEEE J. Sel. Areas Commun. , vol. 40, no. 2, pp. 626–640, Feb. 2021
2021
Earlier work this paper cites.
C. Chen, H. Xu, W. Wang, B. Li, B. Li, L. Chen, and G. Zhang, “Communication-efficient federated learning with adaptive parameter freezing,” in Proc. IEEE Int. Conf. Distrib. Comput. Syst. (ICDCS) , Jul. 2021, pp. 1–11
2021
Earlier work this paper cites.
Y. J. Ha, M. Yoo, S. Park, S. Jung, and J. Kim, “Secure aerial surveillance using split learning,” in Proc. Int. Conf. Ubiquitous Future Netw. (ICUFN) , Jeju Island, Korea, Republic of, Aug. 2021, pp. 434–437
2021
Earlier work this paper cites.
C. Thapa, M. A. P. Chamikara, and S. A. Camtepe, “Advancements of federated learning towards privacy preservation: From federated learning to split learning,” Federated Learning Systems: Towards Next-Generation AI , pp. 79–109, Jun. 2021
2021
Earlier work this paper cites.
J. Shao, Y. Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE J. Sel. Areas Commun. , vol. 40, no. 1, pp. 197–211, Jan. 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. N. Qureshi, M. Manalastas, A. Ijaz, A. Imran, Y. Liu, and M. O. Al Kalaa, “Communication requirements in 5G-enabled healthcare applications: Review and considerations,” in Proc. Healthcare , vol. 10, no. 2, Feb. 2022, p. 293
2022
Earlier work this paper cites.
Y. Guan, Z. Li, Z. Lin, Y. Zhu, J. Leng, and M. Guo, “Block-skim: Efficient question answering for transformer,” in Proc. AAAI Conf. Artif. Intell. (AAAI) , vol. 36, no. 10, Vancouver, BC, Canada, Feb. 2022, pp. 10 710–10 719
2022
Earlier work this paper cites.
A. de Santana Correia and E. L. Colombini, “Attention, please! A survey of neural attention models in deep learning,” Artif. Intell. Rev. , vol. 55, no. 8, pp. 6037–6124, Mar. 2022
2022
Earlier work this paper cites.
W. Yang, Z. Q. Liew, W. Y. B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Cao, and K. B. Letaief, “Semantic communication meets edge intelligence,” IEEE Wireless Commun. , vol. 29, no. 5, pp. 28–35, Dec. 2022
2022
Earlier work this paper cites.
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Tran, Y. Tay, and D. Metzler, “Confident adaptive language modeling,” in Proc. Adv. Neural Inform. Process. Syst. (NeurIPS) , New Orleans, LA, USA, Nov. 2022, pp. 17 456–17 472
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Apr. 2022, pp. 1–13
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Matsubara, M. Levorato, and F. Restuccia, “Split computing and early exiting for deep learning applications: Survey and research challenges,” ACM Comput. Surv. , vol. 55, no. 5, pp. 1–30, Dec. 2022
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” J. Mach. Learn. Res. , vol. 23, no. 120, p. 5232–5270, Jan. 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
X. Chen, G. Zhu, Y. Deng, and Y. Fang, “Federated learning over multihop wireless networks with in-network aggregation,” IEEE Trans. Wireless Commun. , vol. 21, no. 6, pp. 4622–4634, Jun. 2022
2022
Earlier work this paper cites.
X. Liu, K. Ji, Y. Fu, W. L. Tam, Z. Du, Z. Yang, and J. Tang, “P-Tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks,” in Proc. 60th Annu. Meet. Assoc. Comput. Linguist. (ACL) , Dublin, Ireland, May 2022, pp. 61–68
2022
Earlier work this paper cites.
C. Thapa, P. C. M. Arachchige, S. Camtepe, and L. Sun, “SplitFed: When federated learning meets split learning,” in Proc. AAAI Conf. Artif. Intell. (AAAI) , Vancouver, BC, Canada, Feb. 2022, pp. 8485–8493
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Liang, C. Ge, Z. Tong, Y. Song, J. Wang, and P. Xie, “Not all patches are what you need: Expediting vision transformers via token reorganizations,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Apr. 2022, pp. 1–21
2022
Cited alongside, same era.
X. Chen, G. Zhu, H. Ding, L. Zhang, H. Zhang, and Y. Fang, “End-to-end service auction: A general double auction mechanism for edge computing services,” IEEE/ACM Trans. Netw. , vol. 30, no. 6, pp. 2616–2629, Dec. 2022
2022
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Google, “A large language model from Google research, designed for the medical domain,” 2023. [Online]. Available: https://sites.research.google/med-palm/
2023
Cited alongside, same era.
G. DeepMind, “RT-2: New model translates vision and language into action,” 2023. [Online]. Available: https://www.deepmind.com/blog/rt-2-new-model-translates-vision-and-language-into-action
2023
Cited alongside, same era.
M. A. Rahman, “A survey on security and privacy of multimodal LLMs-connected healthcare perspective,” in Proc. IEEE Glob. Commun. Conf. Workshops (GC Wkshps) , Kuala Lumpur, Malaysia, Dec. 2023, pp. 1807–1812
2023
Cited alongside, same era.
Google, “Google introduces Gemini, the most capable and flexible AI model we’ve ever built,” 2023. [Online]. Available: https://store.google.com/intl/en/ideas/articles/pixel-feature-drop-december-2023/
2023
Cited alongside, same era.
Qualcomm, “Qualcomm works with Meta to enable on-device AI applications using Llama 2,” 2023. [Online]. Available: https://www.qualcomm.com/news/releases/2023/07/qualcomm-works-with-meta-to-enable-on-device-ai-applications-usi
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. A. Chien, L. Lin, H. Nguyen, V. Rao, T. Sharma, and R. Wijayawardana, “Reducing the carbon impact of generative AI inference (today and in 2035),” in Proc. 2nd Workshop Sustain. Comput. Syst. , Boston, MA, USA, Jul. 2023, pp. 1–7
2023
Cited alongside, same era.
M. Xu, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, D. I. Kim, and K. B. Letaief, “When large language model agents meet 6G networks: Perception, grounding, and alignment,” IEEE Wireless Commun. , pp. 1–9, early access 2024
2024
Closest in time.
Y. Shen, J. Shao, X. Zhang, Z. Lin, H. Pan, D. Li, J. Zhang, and K. B. Letaief, “Large language models empowered autonomous edge AI for connected intelligence,” IEEE Commun. Mag. , vol. 62, no. 10, pp. 140–146, Oct. 2024
2024
Closest in time.
Y. Huang, H. Du, X. Zhang, D. Niyato, J. Kang, Z. Xiong, S. Wang, and T. Huang, “Large language models for networking: Applications, enabling techniques, and challenges,” IEEE Netw. , pp. 1–8, early access 2024
2024
Closest in time.
H. Zhou, C. Hu, Y. Yuan, Y. Cui, Y. Jin, C. Chen, H. Wu, D. Yuan, L. Jiang, D. Wu et al. , “Large language model (LLM) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,” IEEE Commun. Surveys Tuts. , pp. 1–54, early access 2024
2024
Closest in time.
3GPP, “3rd generation partnership project; Technical specification group services and system aspects; Service requirements for the 5G system; Stage 1 (Release 19),” 3rd Generation Partnership Project (3GPP), Technical Specification (TS) 22.261, Dec. 2024, version 19.6.0
2024
Closest in time.
J. L. Prieto, “New fitbit study explores metabolic health,” 2024. [Online]. Available: https://blog.google/products/fitbit/new-quest-fitbit-study-metabolic-health/
2024
Closest in time.
2024
Closest in time.
Z. Fang, S. Hu, H. An, Y. Zhang, J. Wang, H. Cao, X. Chen, and Y. Fang, “PACP: Priority-aware collaborative perception for connected and autonomous vehicles,” IEEE Trans. Mobile Comput. , pp. 1–15, early access 2024
2024
Closest in time.
X. Chen, Y. Deng, H. Ding, G. Qu, H. Zhang, P. Li, and Y. Fang, “Vehicle as a service (VaaS): Leverage vehicles to build service networks and capabilities for smart cities,” IEEE Commun. Surveys Tuts. , vol. 26, no. 3, pp. 2048–2081, 3rd Quart. 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “PaLM: Scaling language modeling with pathways,” J. Mach. Learn. Res. , vol. 24, no. 1, pp. 11 324–11 436, Mar. 2024
2024
Closest in time.
Claude.ai, “How does Claude 3 AI work?” 2024. [Online]. Available: https://claude3.pro/how-does-claude-3-ai-work/
2024
Closest in time.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol. , vol. 15, no. 3, pp. 1–45, Mar. 2024
2024
Closest in time.
S. Xu, C. K. Thomas, O. Hashash, N. Muralidhar, W. Saad, and N. Ramakrishnan, “Large multi-modal models (LMMs) as universal foundation models for AI-native wireless systems,” IEEE Netw. , vol. 38, no. 5, pp. 10–20, Sep. 2024
2024
Closest in time.
C. Cui, Y. Ma, X. Cao, W. Ye, Y. Zhou, K. Liang, J. Chen, J. Lu, Z. Yang, K.-D. Liao et al. , “A survey on multimodal large language models for autonomous driving,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. , Waikoloa, HI, USA, Jan. 2024, pp. 958–979
2024
Closest in time.
2024
Closest in time.
H. Du, R. Zhang, D. Niyato, J. Kang, Z. Xiong, and D. I. Kim, “Reinforcement learning with large language models (LLMs) interaction for network services,” in Proc. Int. Conf. Comput. Netw. Commun. (ICNC) , Feb. 2024, pp. 799–803
2024
Closest in time.
2024
Closest in time.
R. Zhang, H. Du, Y. Liu, D. Niyato, J. Kang, S. Sun, X. Shen, and H. V. Poor, “Interactive AI with retrieval-augmented generation for next generation networking,” IEEE Netw. , pp. 1–11, early access 2024
2024
Closest in time.
Y. Zhu, Y. Liu, F. Stahlberg, S. Kumar, Y.-h. Chen, L. Luo, L. Shu, R. Liu, J. Chen, and L. Meng, “Towards an on-device agent for text rewriting,” in Proc. North Am. Chapter Assoc. Comput. Linguist. (NAACL) , Mexico City, Mexico, Jun. 2024, pp. 2535–2552
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Chen, Z. Guo, W. Meng, S. Han, C. Li, and T. Q. Quek, “A survey on resource management in joint communication and computing-embedded SAGIN,” IEEE Commun. Surveys Tuts. , pp. 1–44, early access 2024
2024
Closest in time.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “AWQ: Activation-aware weight quantization for LLM compression and acceleration,” in Proc. Mach. Learn. Syst. , vol. 6, Santa Clara, CA, USA, May 2024, pp. 87–100
2024
Closest in time.
Y. Gu, L. Dong, F. Wei, and M. Huang, “MiniLLM: Knowledge distillation of large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–24
2024
Closest in time.
M. Wu, A. Waheed, C. Zhang, M. Abdul-Mageed, and A. F. Aji, “LaMini-LM: A diverse herd of distilled models from large-scale instructions,” in Proc. Conf. Eur. Chapter Assoc. Comput. Linguist. (EACL) , St. Julian’s, Malta, Mar. 2024, pp. 944–964
2024
Closest in time.
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao, “Medusa: Simple LLM inference acceleration framework with multiple decoding heads,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, 2024, pp. 1–27
2024
Closest in time.
Y. Fu, P. Bailis, I. Stoica, and H. Zhang, “Break the sequential dependency of LLM inference using lookahead decoding,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, Jul. 2024, pp. 1–20
2024
Closest in time.
2024
Closest in time.
Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V. Braverman, B. Chen, and X. Hu, “KIVI: A tuning-free asymmetric 2bit quantization for KV cache,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, Jul. 2024, pp. 32 332–32 344
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y. Qiao, and P. Luo, “OmniQuant: Omnidirectionally calibrated quantization for large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–25
2024
Closest in time.
2024
Closest in time.
G. Gerganov, “llama.cpp,” 2024, accessed: 2024-05-10. [Online]. Available: https://github.com/ggerganov/llama.cpp
2024
Closest in time.
Y. Chen, X. Pan, Y. Li, B. Ding, and J. Zhou, “EE-LLM: Large-scale training and inference of early-exit large language models with 3D parallelism,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, jul 2024, pp. 1–27
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Ribar, I. Chelombiev, L. Hudlass-Galley, C. Blake, C. Luschi, and D. Orr, “SparQ attention: Bandwidth-efficient LLM inference,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, jul 2024, pp. 1–26
2024
Closest in time.
L. Song, Y. Chen, S. Yang, X. Ding, Y. Ge, Y.-C. Chen, and Y. Shan, “Low-rank approximation for sparse attention in multi-modal llms,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , Seattle, WA, USA, Jun. 2024, pp. 13 763–13 773
2024
Closest in time.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–21
2024
Closest in time.
WWDC24, “Platforms state of the union (ASL),” 2024. [Online]. Available: https://developer.apple.com/videos/play/wwdc2024/102/?time=279
2024
Closest in time.
C. Wu, J. Chen, S. Zhu, W. Feng, K. He, R. Du, and Y. Xiang, “WAFBOOSTER: Automatic boosting of waf security against mutated malicious payloads,” IEEE Trans. Depend. Sec. Comput. , pp. 1–13, early access 2024
2024
Closest in time.
2024
Closest in time.
G. Qu, Z. Lin, F. Liu, X. Chen, and K. Huang, “TrimCaching: Parameter-sharing AI model caching in wireless edge networks,” in Proc. IEEE Int. Conf. Distrib. Comput. Syst. (ICDCS) , Jersey City, NJ, USA, Jul. 2024, pp. 36–46
2024
Closest in time.
H. Xiong, J. Bian, Y. Li, X. Li, M. Du, S. Wang, D. Yin, and S. Helal, “When search engine services meet large language models: Visions and challenges,” IEEE Trans. Serv. Comput. , pp. 1–23, early access 2024
2024
Closest in time.
P. Ghimire, K. Kim, and M. Acharya, “Opportunities and challenges of generative AI in construction industry: Focusing on adoption of text-based models,” Buildings , vol. 14, no. 1, p. 220, Jan. 2024
2024
Closest in time.
2024
Closest in time.
Z. Lin, G. Qu, X. Chen, and K. Huang, “Split learning in 6G edge networks,” IEEE Wireless Commun. , vol. 31, no. 4, pp. 170–176, Aug. 2024
2024
Closest in time.
W. Fan, Y. Ding, L. Ning, S. Wang, H. Li, D. Yin, T.-S. Chua, and Q. Li, “A survey on RAG meeting LLMs: Towards retrieval-augmented large language models,” in Proc. ACM SIGKDD Conf. Knowl. Discov. Data Min. , Barcelona Spain, Aug. 2024, pp. 6491–6501
2024
Closest in time.
J. Chen, H. Lin, X. Han, and L. Sun, “Benchmarking large language models in retrieval-augmented generation,” in Proc. AAAI Conf. Artif. Intell. (AAAI) , vol. 38, no. 16, Vancouver, BC, Canada, Feb. 2024, pp. 17 754–17 762
2024
Closest in time.
2024
Closest in time.
H. Wu, Q. Zeng, and K. Huang, “Efficient multiuser AI downloading via reusable knowledge broadcasting,” IEEE Trans. Wireless Commun. , vol. 23, no. 8, pp. 10 459–10 472, Aug. 2024
2024
Closest in time.
2024
Closest in time.
S. Kim, C. Hooper, A. Gholami, Z. Dong, X. Li, S. Shen, M. W. Mahoney, and K. Keutzer, “SqueezeLLM: Dense-and-sparse quantization,” in Proc. Int. Conf. Mach. Learn. (ICML) , Vienna, Austria, jul 2024, pp. 1–23
2024
Closest in time.
T. Dettmers, R. Svirschevski, V. Egiazarian, D. Kuznedelev, E. Frantar, S. Ashkboos, A. Borzunov, T. Hoefler, and D. Alistarh, “SpQR: A sparse-quantized representation for near-lossless LLM weight compression,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–29
2024
Closest in time.
J. Jiang, X. Liu, and C. Fan, “Low-parameter federated learning with large language models,” in International Conference on Web Information Systems and Applications , Yinchuan, China, Aug. 2024, pp. 319–330
2024
Closest in time.
W. Kuang, B. Qian, Z. Li, D. Chen, D. Gao, X. Pan, Y. Xie, Y. Li, B. Ding, and J. Zhou, “FederatedScope-LLM: A comprehensive package for fine-tuning large language models in federated learning,” in Proc. ACM SIGKDD Conf. Knowl. Discovery Data Mining (KDD) , Barcelona, Spain, Aug. 2024, p. 5260–5271
2024
Closest in time.
G. Wang, J. Liu, C. Li, J. Ma, Y. Zhang, X. Wei, K. Zhang, M. Chong, R. Zhang, Y. Liu, and S. Zhang, “Cloud-device collaborative learning for multimodal large language models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , Seattle, WA, USA, Jun. 2024, pp. 12 646–12 655
2024
Closest in time.
2024
Closest in time.
A. M. Devices, “Fine-tune Llama 2 with LoRA: Customizing a large language model for question-answering,” 2024. [Online]. Available: https://rocm.blogs.amd.com/artificial-intelligence/llama2-lora/README.html
2024
Closest in time.
Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong et al. , “MegaScale: Scaling large language model training to more than 10,000 GPUs,” in Proc. USENIX Symp. Netw. Syst. Des. Implement. (NSDI) , Santa Clara, CA, USA, Apr. 2024, pp. 745–760
2024
Closest in time.
J. Zhou, Y. Chen, Z. Hong, W. Chen, Y. Yu, T. Zhang, H. Wang, C. Zhang, and Z. Zheng, “Training and serving system of foundation models: A comprehensive survey,” IEEE Open J. Comput. Soc. , vol. 5, pp. 107–119, Mar. 2024
2024
Closest in time.
2024
Closest in time.
Z. Lin, Z. Chen, Z. Fang, X. Chen, X. Wang, and Y. Gao, “FedSN: A general federated learning framework over LEO satellite networks,” IEEE Trans. Mobile Comput. , pp. 1–15, early access 2024
2024
Closest in time.
2024
Closest in time.
H. Woisetschläger, A. Isenko, S. Wang, R. Mayer, and H.-A. Jacobsen, “Federated fine-tuning of LLMs on the very edge: The good, the bad, the ugly,” in Proc. Workshop Data Manag. End-to-End Mach. Learn. , Santiago, AA, Chile, Jun. 2024, p. 39–50
2024
Closest in time.
J. Zhang, S. Vahidian, M. Kuo, C. Li, R. Zhang, T. Yu, Y. Zhou, G. Wang, and Y. Chen, “Towards building the federated GPT: Federated instruction tuning,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Seoul, Korea, Republic of, Apr. 2024, pp. 6915–6919
2024
Closest in time.
Z. Lin, G. Zhu, Y. Deng, X. Chen, Y. Gao, K. Huang, and Y. Fang, “Efficient parallel split learning over resource-constrained wireless edge networks,” IEEE Trans. Mobile Comput. , vol. 23, no. 10, pp. 9224–9239, Oct. 2024
2024
Closest in time.
2024
Closest in time.
G. Zhu, Y. Deng, X. Chen, H. Zhang, Y. Fang, and T. F. Wong, “ESFL: Efficient split federated learning over resource-constrained heterogeneous wireless devices,” IEEE Internet Things J. , vol. 11, no. 16, pp. 27 153–27 166, Aug. 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. Ruhle, L. V. Lakshmanan, and A. H. Awadallah, “Hybrid LLM: Cost-efficient and quality-aware query routing,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–19
2024
Closest in time.
2024
Closest in time.
Q. Cao, S. Min, Y. Wang, and H. Hajishirzi, “BTR: Binary token representations for efficient retrieval augmented language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–18
2024
Closest in time.
2024
Closest in time.
J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang, “QLLM: Accurate and efficient low-bitwidth quantization for large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–23
2024
Closest in time.
Y. Jin, K. Xu, L. Chen, C. Liao, J. Tan, B. Chen, C. Lei, A. Liu, C. Song, X. Lei et al. , “Unified language-vision pretraining in LLM with dynamic discrete visual tokenization,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–21
2024
Closest in time.
P. Jin, R. Takanobu, C. Zhang, X. Cao, and L. Yuan, “Chat-UniVi: Unified visual representation empowers large language models with image and video understanding,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , Seattle, WA, USA, Jun. 2024, pp. 13 700–13 710
2024
Closest in time.
J. Huang, D. Li, C. Huang, X. Qin, and W. Zhang, “Joint task and data-oriented semantic communications: A deep separate source-channel coding scheme,” IEEE Internet Things J. , vol. 11, no. 2, pp. 2255–2272, Jan. 2024
2024
Closest in time.
P. Wang, J. Li, C. Liu, X. Fan, M. Ma, and Y. Wang, “Distributed semantic communications for multimodal audio-visual parsing tasks,” IEEE Trans. Green Commun. Netw. , pp. 1–10, early access, 2024
2024
Closest in time.
J. Huang, K. Yuan, C. Huang, and K. Huang, “D 2
2024
Closest in time.
M. Aibin, “Energy consumption of ChatGPT responses,” 2024. [Online]. Available: https://www.baeldung.com/cs/chatgpt-large-language-models-power-consumption
2024
Closest in time.
2024
Closest in time.
Y. Mao, X. Yu, K. Huang, Y.-J. A. Zhang, and J. Zhang, “Green edge AI: A contemporary survey,” Proc. IEEE , pp. 1–32, early access 2024
2024
Closest in time.
C. Wu, J. Chen, Q. Fang, K. He, Z. Zhao, H. Ren, G. Xu, Y. Liu, and Y. Xiang, “Rethinking membership inference attacks against transfer learning,” IEEE Trans. Inf. Forensics Secur. , vol. 19, pp. 6441–6454, Jun. 2024
2024
Closest in time.
A. Panda, C. A. Choquette-Choo, Z. Zhang, Y. Yang, and P. Mittal, “Teach LLMs to phish: Stealing private information from language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vienna Austria, May 2024, pp. 1–25
2024
Closest in time.
M. Hofer, D. Obraczka, A. Saeedi, H. Köpcke, and E. Rahm, “Construction of knowledge graphs: Current state and challenges,” Information , vol. 15, no. 8, p. 509, Aug. 2024
2024
Closest in time.
Wonderchat, “How to train ChatGPT on your own data,” 2024. [Online]. Available: https://wonderchat.io/blog/how-to-train-chatgpt-on-your-own-data
2024
Closest in time.
J. Zhang, H. Gao, P. Zhang, B. Feng, W. Deng, and Y. Hou, “LA-UCL: LLM-augmented unsupervised contrastive learning framework for few-shot text classification,” in Proc. Joint Int. Conf. Comput. Linguistics, Lang. Resources Eval. (LREC-COLING 2024) , Torino, Italia, May 2024, pp. 10 198–10 207
2024
Closest in time.
HuaweiTech, “ITU-R WP5D completed the recommendation framework for IMT-2030 (global 6G vision),” 2023. [Online]. Available: https://www.huawei.com/en/huaweitech/future-technologies/itu-r-wp5d-completed-recommendation-framework-imt-2030
2030
Closest in time.
R. Tandon and O. Simeone, “Cloud-aided wireless networks with edge caching: Fundamental latency trade-offs in fog radio access networks,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT) , Barcelona, Spain, Jul. 2016, pp. 2029–2033
2033
Closest in time.