Fetching the paper…
Reading the bibliography…
Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities.
Y. Hu, H. Jiang, D. Feng, L. Tian, S. Zhang, J. Liu, W. Tong, Y. Qin, and L. Wang, “Achieving page-mapping ftl performance at block-mapping ftl cost by hiding address translation,” in 2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST) , 2010
2010
Earlier work this paper cites.
Y. Cai, G. Yalcin, O. Mutlu, E. F. Haratsch, A. Crista, O. S. Unsal, and K. Mai, “Error analysis and retention-aware error management for nand flash memory.” Intel Technology Journal , vol. 17, no. 1, 2013
2013
Earlier work this paper cites.
J. Do, Y.-S. Kee, J. M. Patel, C. Park, K. Park, and D. J. DeWitt, “Query processing on smart ssds: Opportunities and challenges,” in Proceedings of the 2013 ACM SIGMOD International Conference on Management of Data , 2013
2013
Earlier work this paper cites.
D. Pandiyan and C.-J. Wu, “Quantifying the energy cost of data movement for emerging smart phone workloads on mobile platforms,” in 2014 IEEE International Symposium on Workload Characterization (IISWC) , 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” 2015
2015
Earlier work this paper cites.
Y. Cai, S. Ghose, E. F. Haratsch, Y. Luo, and O. Mutlu, “Error characterization, mitigation, and recovery in flash-memory-based solid-state drives,” Proceedings of the IEEE , vol. 105, no. 9, pp. 1666–1704, 2017
2017
Earlier work this paper cites.
G. Koo, K. K. Matam, T. I, H. K. G. Narra, J. Li, H.-W. Tseng, S. Swanson, and M. Annavaram, “Summarizer: trading communication with computing near storage,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. K. Gonugondla, M. Kang, Y. Kim, M. Helm, S. Eilert, and N. Shanbhag, “Energy-efficient deep in-memory architecture for nand flash memories,” in 2018 IEEE International Symposium on Circuits and Systems (ISCAS) , 2018
2018
Earlier work this paper cites.
S.-W. Jun, A. Wright, S. Zhang, S. Xu et al. , “Grafboost: Using accelerated flash storage for external graph analytics,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) , 2018
2018
Earlier work this paper cites.
M. Brysbaert, “How many words do we read per minute? a review and meta-analysis of reading rate,” Journal of memory and language , vol. 109, p. 104047, 2019
2019
Earlier work this paper cites.
V. S. Mailthody, Z. Qureshi, W. Liang, Z. Feng, S. G. De Gonzalo, Y. Li, H. Franke, J. Xiong, J. Huang, and W.-m. Hwu, “Deepstore: In-storage acceleration for intelligent queries,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019
2019
Earlier work this paper cites.
K. K. Matam, G. Koo, H. Zha, H.-W. Tseng, and M. Annavaram, “Graphssd: graph semantics aware ssd,” in Proceedings of the 46th international symposium on computer architecture , 2019
2019
Earlier work this paper cites.
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Processing data where it makes sense: Enabling in-memory computation,” Microprocessors and Microsystems , vol. 67, pp. 28–41, 2019
2019
Earlier work this paper cites.
M. Naumov, D. Mudigere, H.-J. M. Shi, J. Huang, N. Sundaraman, J. Park, X. Wang, U. Gupta, C.-J. Wu, A. G. Azzolini, D. Dzhulgakov, A. Mallevich, I. Cherniavskii, Y. Lu, R. Krishnamoorthi, A. Yu, V. Kondratenko, S. Pereira, X. Chen, W. Chen, V. Rao, B. Jia, L. Xiong, and M. Smelyanskiy, “Deep learning recommendation model for personalization and recommendation systems,” 2019
2019
Earlier work this paper cites.
M. Torabzadehkashi, S. Rezaei, A. Heydarigorji, H. Bobarshad, V. Alves, and N. Bagherzadeh, “Catalina: In-storage processing acceleration for scalable big data analytics,” in 2019 27th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP) , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Gupta, J. Morris, M. Imani, R. Ramkumar, J. Yu, A. Tiwari, B. Aksanli, and T. Š. Rosing, “Thrifty: Training with hyperdimensional computing across flash hierarchy,” in Proceedings of the 39th International Conference on Computer-Aided Design , 2020
2020
Earlier work this paper cites.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J.-H. Park, S. Lee, K. Park, J. W. Lee et al. , “Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020
2020
Earlier work this paper cites.
B. Li, S. Pandey, H. Fang, Y. Lyv, J. Li, J. Chen, M. Xie, L. Wan, H. Liu, and C. Ding, “Ftrans: energy-efficient acceleration of transformers using fpga,” in Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design , 2020
2020
Earlier work this paper cites.
S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC) , 2020
2020
Earlier work this paper cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020
2020
Earlier work this paper cites.
A. Sebastian, M. Le Gallo, R. Khaddam-Aljameh, and E. Eleftheriou, “Memory devices and applications for in-memory computing,” Nature nanotechnology , vol. 15, no. 7, pp. 529–544, 2020
2020
Earlier work this paper cites.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed, “Big bird: Transformers for longer sequences,” in Advances in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
S. Kim, Y. Jin, G. Sohn, J. Bae, T. J. Ham, and J. W. Lee, “Behemoth: a flash-centric training accelerator for extreme-scale { \{ DNNs } \} ,” in 19th USENIX Conference on File and Storage Technologies (FAST 21) , 2021
2021
Cited alongside, same era.
J. Kosaian and K. Rashmi, “Arithmetic-intensity-guided fault tolerance for neural network inference on gpus,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Cited alongside, same era.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , 2021
2021
Cited alongside, same era.
H. Peng, S. Huang, T. Geng, A. Li, W. Jiang, H. Liu, S. Wang, and C. Ding, “Accelerating transformer-based deep learning models on fpgas using column balanced block pruning,” in 2021 22nd International Symposium on Quality Electronic Design (ISQED) , 2021
P. Hämäläinen, M. Tavast, and A. Kunnari, “Evaluating large language models in generating synthetic hci research data: a case study,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023
2023
Later among the works it cites.
Y. Jin, C.-F. Wu, D. Brooks, and G.-Y. Wei, “Sˆ3: Increasing gpu utilization during generative inference for higher throughput,” in Advances in Neural Information Processing Systems , 2023
2023
Later among the works it cites.
B. Kim, S. Lee, B. Hah, K. Park, Y. Park, K. Jo, Y. Noh, H. Seol, H. Lee, J. Shin et al. , “28.2 a high-performance 1tb 3b/cell 3d-nand flash with a 194mb/s write throughput on over 300 layers,” in 2023 IEEE International Solid-State Circuits Conference (ISSCC) , 2023
2023
Later among the works it cites.
J. Kim, M. Kang, Y. Han, Y.-G. Kim, and L.-S. Kim, “Optimstore: In-storage optimization of large scale dnns with on-die processing,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
P. Qi, Y. Song, H. Peng, S. Huang, Q. Zhuge, and E. H.-M. Sha, “Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization,” in Proceedings of the 2021 on Great Lakes Symposium on VLSI , 2021
2021
Cited alongside, same era.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM , vol. 64, no. 9, pp. 99–106, 2021
2021
Cited alongside, same era.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2021
2021
Cited alongside, same era.
M. Wilkening, U. Gupta, S. Hsia, C. Trippel, C.-J. Wu, D. Brooks, and G.-Y. Wei, “Recssd: near data processing for solid state drive based recommendation inference,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2021
2021
Cited alongside, same era.
X. Zhang, Y. Wu, P. Zhou, X. Tang, and J. Hu, “Algorithm-hardware co-design of attention mechanism on fpga devices,” ACM Transactions on Embedded Computing Systems (TECS) , vol. 20, no. 5s, pp. 1–24, 2021
2021
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel, “Palm: Scaling language modeling with pathways,” 2022
2022
Cited alongside, same era.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” in Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Llm.int8(): 8-bit matrix multiplication for transformers at scale,” ArXiv , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Kim, C. Jung, S. Yoo, D. Hong, J. Hwang, J. Yoon, O. Jung, J. Choi, S. Hyun, M. Kang et al. , “A 1.1 v 16gb ddr5 dram with probabilistic-aggressor tracking, refresh-management functionality, per-row hammer tracking, a multi-step precharge, and core-bias modulation for security and reliability enhancement,” in 2023 IEEE International Solid-State Circuits Conference (ISSCC) , 2023
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023
2023
Later among the works it cites.
S. Li, F. Tu, L. Liu, J. Lin, Z. Wang, Y. Kang, Y. Ding, and Y. Xie, “Ecssd: Hardware/data layout co-designed in-storage-computing architecture for extreme classification,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Re, I. Stoica, and C. Zhang, “FlexGen: High-throughput generative inference of large language models with a single GPU,” in Proceedings of the 40th International Conference on Machine Learning , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/mlc-ai/mlc-llm
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “SmoothQuant: Accurate and efficient post-training quantization for large language models,” in Proceedings of the 40th International Conference on Machine Learning , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Ye, X. Zhou, J. Zhou, C. Chen, and K. Li, “Accelerating attention mechanism on fpgas based on efficient reconfigurable systolic array,” ACM Transactions on Embedded Computing Systems , vol. 22, no. 6, pp. 1–22, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Zhao, S. Yang, K. Xie, Y. Feng, Q. Wang, P. Sang, X. Zhan, J. Wu, and J. Chen, “Error bits recovering in 3d NAND flash memory: A novel state-shift re-program (SRP) scheme,” in IEEE International Conference on Integrated Circuits, Technologies and Applications, ICTA 2023, Hefei, China, October 27-29, 2023 , 2023
2023
Later among the works it cites.
K. Alizadeh, I. Mirzadeh, D. Belenko, K. Khatamifard, M. Cho, C. C. D. Mundo, M. Rastegari, and M. Farajtabar, “Llm in a flash: Efficient large language model inference with limited memory,” 2024
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
W. Jung, H. Kim, D.-B. Kim, T.-H. Kim, N. Lee, D. Shin, M. Kim, Y. Rho, H.-J. Lee, Y. Hyun et al. , “13.3 a 280-layer 1tb 4b/cell 3d-nand flash memory with a 28.5 gb/mm2 areal density and a 3.2 gb/s high-speed io rate,” in 2024 IEEE International Solid-State Circuits Conference (ISSCC) , 2024
2024
Closest in time.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for llm compression and acceleration,” 2024
2024
Closest in time.
Y. Seo, J. Choi, S. Cho, H. Han, W. Kim, G. Ryu, J. Ahn, Y. Cho, S. Choi, S. Lee et al. , “13.8 a 1a-nm 1.05 v 10.5 gb/s/pin 16gb lpddr5 turbo dram with wck correction strategy, a voltage-offset-calibrated receiver and parasitic capacitance reduction,” in 2024 IEEE International Solid-State Circuits Conference (ISSCC) , 2024
2024
Closest in time.
Y. Wang, X. Pan, Y. An, J. Zhang, and G. Reinman, “Beacongnn: Large-scale gnn acceleration with out-of-order streaming in-storage computing,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2024, pp. 330–344
2024
Closest in time.