Fetching the paper…
Reading the bibliography…
By provisioning inference offloading services, edge inference drives the rapid growth of AI applications at network edge.
P. Brucker and M. Y. Kovalyov, “Single machine batch scheduling to minimize the weighted number of late jobs,” Math. Methods Oper. Res. , vol. 43, pp. 1–8, Feb. 1996
1996
Earlier work this paper cites.
A. Allahverdi, C. T. Ng, T. E. Cheng, and M. Y. Kovalyov, “A survey of scheduling problems with setup times or costs,” Eur. J. Oper. Res. , vol. 187, no. 3, pp. 985–1032, Jun. 2008
2008
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” Apr. 2009
2009
Earlier work this paper cites.
R. Love, Linux Kernel Development (3rd Edition) , 3rd ed., ser. Developer’s Library. Addison-Wesley, 2010
2010
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Proc. Int. Conf. Neural Inf. Process. Syst. , Montreal, Canada, Dec. 2014, pp. 3320–3328
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) , Las Vegas, NV, USA, Jun. 2016, pp. 770–778
2016
Earlier work this paper cites.
G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Vancouver, Cananda, Apr. 2018, pp. 1–18
2018
Earlier work this paper cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in Proc. 56th Annu. Meet. Assoc. Comput. Linguistics (ACL) , Melbourne, Australia, Jul. 2018, pp. 328–339
2018
Earlier work this paper cites.
J. Zhang and K. B. Letaief, “Mobile edge intelligence and computing for the internet of vehicles,” Proc. IEEE , vol. 108, no. 2, pp. 246–261, Feb. 2019
2019
Earlier work this paper cites.
W. Dai, H. Nishi, V. Vyatkin, V. Huang, Y. Shi, and X. Guan, “Industrial edge computing: Enabling embedded intelligence,” IEEE Ind. Electron. Mag. , vol. 13, no. 4, pp. 48–56, Dec. 2019
2019
Earlier work this paper cites.
Y. Guo, H. Shi, A. Kumar, K. Grauman, T. Rosing, and R. Feris, “Spottune: Transfer learning through adaptive fine-tuning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) , Long Beach, CA, USA, Jun. 2019
2019
Earlier work this paper cites.
Q. Yang, X. Luo, P. Li, T. Miyazaki, and X. Wang, “Computation offloading for fast CNN inference in edge computing,” in Proc. Conf. Res. Adapt. Converg. Syst. , Chongqing China, Sep. 2019, pp. 101–106
2019
Earlier work this paper cites.
J. Ren, G. Yu, Y. He, and G. Y. Li, “Collaborative cloud and edge computing for latency minimization,” IEEE Trans. Veh. Technol. , vol. 68, no. 5, pp. 5031–5044, May 2019
2019
Earlier work this paper cites.
S. U. Amin and M. S. Hossain, “Edge intelligence and internet of things in healthcare: A survey,” IEEE Access , vol. 9, pp. 45–59, Dec. 2020
2020
Earlier work this paper cites.
Y. Shi, K. Yang, T. Jiang, J. Zhang, and K. B. Letaief, “Communication-efficient edge AI: Algorithms and systems,” IEEE Commun. Surveys Tuts. , vol. 22, no. 4, pp. 2167–2191, 4th Quart. 2020
2020
Earlier work this paper cites.
J. Shao and J. Zhang, “Communication-computation trade-off in resource-constrained edge inference,” IEEE Commun. Mag. , vol. 58, no. 12, pp. 20–26, Dec. 2020
2020
Earlier work this paper cites.
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y. Zhu, H. Zhu, H. Xiong, and Q. He, “A comprehensive survey on transfer learning,” Proc. IEEE , vol. 109, no. 1, pp. 43–76, Jan. 2020
2020
Cited alongside, same era.
NVIDIA, “Nvidia AMPERE GA102 GPU architecture,” 2020. [Online]. Available: https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf
2020
Cited alongside, same era.
3GPP, “3rd generation partnership project; Technical specification group services and system aspects; Study on traffic characteristics and performance requirements for AI/ML model transfer in 5GS; (Release 18),” 3rd Generation Partnership Project (3GPP), Technical Specification (TS) 22.874, Dec. 2021, version 18.2.0
2021
Cited alongside, same era.
S. S. Ogden, G. R. Gilman, R. J. Walls, and T. Guo, “Many models at the edge: Scaling deep inference via model-level caching,” in Proc. IEEE Int. Conf. Auton. Comput. Self-Organizing Syst. (ACSOS) , Washington, DC, USA, Sep. 2021, pp. 51–60
2021
F. Xu, J. Xu, J. Chen, L. Chen, R. Shang, Z. Zhou, and F. Liu, “iGniter: Interference-aware GPU resource provisioning for predictable DNN inference in the cloud,” IEEE Trans. Parallel Distrib. Syst. , vol. 34, no. 3, pp. 812–827, Mar. 2023
2023
Later among the works it cites.
N. Ding, Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.-M. Chan, W. Chen et al. , “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nat. Mach. Intell. , vol. 5, no. 3, pp. 220–235, Mar. 2023
2023
Later among the works it cites.
Z. Fu, H. Yang, A. M.-C. So, W. Lam, L. Bing, and N. Collier, “On the effectiveness of parameter-efficient fine-tuning,” in Proc. AAAI Conf. Artif. Intell. (AAAI) , Washington, DC, USA, Feb. 2023, pp. 12 799–12 807
2023
Later among the works it cites.
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Li, J. Gao, L. Zhao, and X. Shen, “Adaptive computing scheduling for edge-assisted autonomous driving,” IEEE Trans. Veh. Technol. , vol. 70, no. 6, pp. 5318–5331, Jun. 2021
2021
Cited alongside, same era.
Z. Yang, K. Nahrstedt, H. Guo, and Q. Zhou, “DeepRT: A soft real time scheduler for computer vision applications on the edge,” in Proc. IEEE/ACM Symp. Edge Comput. (SEC) , San Jose, CA, USA, Dec. 2021, pp. 271–284
2021
Cited alongside, same era.
R. He, L. Liu, H. Ye, Q. Tan, B. Ding, L. Cheng, J. Low, L. Bing, and L. Si, “On the effectiveness of adapter-based tuning for pretrained language model adaptation,” in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics and 11th Int. Joint Conf. Natural Lang. Process. , Aug. 2021, pp. 2208–2222
2021
Cited alongside, same era.
W. Y. B. Lim, Z. Xiong, D. Niyato, X. Cao, C. Miao, S. Sun, and Q. Yang, “Realizing the metaverse with edge intelligence: A match made in heaven,” IEEE Wireless Commun. , vol. 30, no. 4, pp. 64–71, Aug. 2022
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , Apr. 2022, pp. 1–13
2022
Cited alongside, same era.
W. Shi, S. Zhou, Z. Niu, M. Jiang, and L. Geng, “Multiuser co-inference with batch processing capable edge server,” IEEE Trans. Wireless Commun. , vol. 22, no. 1, pp. 286–300, Jan. 2022
2022
Cited alongside, same era.
M. Yao, L. Chen, J. Zhang, J. Huang, and J. Wu, “Loading cost-aware model caching and request routing for cooperative edge inference,” in Proc. IEEE Int. Conf. Commun. (ICC) , Seoul, South Korea, May 2022, pp. 2327–2332
2022
Cited alongside, same era.
Y. D. Kwon, J. Chauhan, and C. Mascolo, “YONO: Modeling multiple heterogeneous neural networks on microcontrollers,” in Proc. ACM/IEEE Int. Conf. Inf. Process. Sensor Netw. (IPSN) , Milano, Italy, May 2022, pp. 285–297
2022
Cited alongside, same era.
Later among the works it cites.
X. Han, Z. Cai, Y. Zhang, C. Fan, J. Liu, R. Ma, and R. Buyya, “Hermes: Memory-efficient pipeline inference for large models on edge devices,” in Proc. IEEE Int. Conf. Computer Design (ICCD) , Milan, Italy, Nov. 2024, pp. 454–461
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Cang, M. Chen, and K. Huang, “Joint batching and scheduling for high-throughput multiuser edge AI with asynchronous task arrivals,” IEEE Trans. Wireless Commun. , vol. 23, no. 10, pp. 13 782–13 795, Oct. 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Ma, K. Gao, and C. Zhang, “Efficient latency optimization in batch processing server systems for multi-user co-inference,” in Proc. IEEE/CIC Int. Conf. Commun. China Workshops (ICCC Workshops) , Hangzhou, China, Aug. 2024, pp. 475–480
2024
Later among the works it cites.
Y. She, T. Shi, J. Wang, and B. Liu, “Dynamic batching and early-exiting for accurate and timely edge inference,” in Proc. IEEE Veh. Technol. Conf. (VTC2024-Spring) , Singapore, Jun. 2024, pp. 1–6
2024
Later among the works it cites.
Y. She, T. Shi, J. Wang, and B. Liu, “Dynamic batching and early-exiting for accurate and timely edge inference,” in Proc. IEEE Veh. Technol. Conf. (VTC) , Singapore, Singapore, Jun. 2024, pp. 1–6
2024
Later among the works it cites.
Z. Zhang, Y. Zhao, H. Li, and J. Liu, “BCEdge: SLO-aware DNN inference services with adaptive batch-concurrent scheduling on edge devices,” IEEE Trans. Netw. Service Manag. , vol. 21, no. 4, pp. 4131–4145, Aug. 2024
2024
Later among the works it cites.
H. Wang, T. Li, M. Zhang, Q. Li, H. Cui, Y. Jiang, and Z. Yuan, “Joint configuration optimization and GPU allocation for multi-tenant real-time video analytics on resource-constrained edge,” IEEE Trans. Mobile Comput. , early access 2024
2024
Later among the works it cites.
J. Wu, L. Wang, Q. Jin, and F. Liu, “Graft: Efficient inference serving for hybrid deep learning with SLO guarantees via DNN re-alignment,” IEEE Trans. Parallel Distrib. Syst. , vol. 35, no. 2, pp. 280–296, Feb. 2024
2024
Later among the works it cites.
G. Qu, Q. Chen, W. Wei, Z. Lin, X. Chen, and K. Huang, “Mobile edge intelligence for large language models: A contemporary survey,” IEEE Commun. Surveys Tuts. , pp. 1–42, early access 2025
2025
Closest in time.