Fetching the paper…
Reading the bibliography…
The widespread of Large Language Models (LLMs) marks a significant milestone in generative AI.
N. Agrawal, V. Prabhakaran, T. Wobber, J. D. Davis, M. Manasse, and R. Panigrahy, “Design tradeoffs for { \{ SSD } \} performance,” in 2008 USENIX Annual Technical Conference (USENIX ATC 08) , 2008
2008
Earlier work this paper cites.
N. Agrawal, V. Prabhakaran, T. Wobber, J. D. Davis, M. Manasse, and R. Panigrahy, “Design tradeoffs for { \{ SSD } \} performance,” in 2008 USENIX Annual Technical Conference (USENIX ATC 08) , 2008
2008
Earlier work this paper cites.
X.-Y. Hu, E. Eleftheriou, R. Haas, I. Iliadis, and R. Pletka, “Write amplification analysis in flash-based solid state drives,” in Proceedings of SYSTOR 2009: The Israeli Experimental Systems Conference , 2009, pp. 1–9
2009
Earlier work this paper cites.
X.-Y. Hu, E. Eleftheriou, R. Haas, I. Iliadis, and R. Pletka, “Write amplification analysis in flash-based solid state drives,” in Proceedings of SYSTOR 2009: The Israeli Experimental Systems Conference , 2009, pp. 1–9
2009
Earlier work this paper cites.
P. Desnoyers, “Analytic modeling of ssd write performance,” in Proceedings of the 5th Annual International Systems and Storage Conference , 2012, pp. 1–10
2012
Earlier work this paper cites.
P. Desnoyers, “Analytic modeling of ssd write performance,” in Proceedings of the 5th Annual International Systems and Storage Conference , 2012, pp. 1–10
2012
Earlier work this paper cites.
J.-W. Hsieh, H.-Y. Lin, and D.-L. Yang, “Multi-channel architecture-based ftl for reliable and high-performance ssd,” IEEE Transactions on Computers , vol. 63, no. 12, pp. 3079–3091, 2013
2013
Earlier work this paper cites.
J.-W. Hsieh, H.-Y. Lin, and D.-L. Yang, “Multi-channel architecture-based ftl for reliable and high-performance ssd,” IEEE Transactions on Computers , vol. 63, no. 12, pp. 3079–3091, 2013
2013
Earlier work this paper cites.
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, “Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics . Vancouver, Canada: Association for Computational Linguistics, July 2017
2017
Earlier work this paper cites.
G. Koo, K. K. Matam, T. I, H. K. G. Narra, J. Li, H.-W. Tseng, S. Swanson, and M. Annavaram, “Summarizer: trading communication with computing near storage,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , 2017, pp. 219–231
2017
Earlier work this paper cites.
B. Mao, S. Wu, and L. Duan, “Improving the ssd performance by exploiting request characteristics and internal parallelism,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 37, no. 2, pp. 472–484, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, “Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics . Vancouver, Canada: Association for Computational Linguistics, July 2017
2017
Earlier work this paper cites.
G. Koo, K. K. Matam, T. I, H. K. G. Narra, J. Li, H.-W. Tseng, S. Swanson, and M. Annavaram, “Summarizer: trading communication with computing near storage,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , 2017, pp. 219–231
2017
Earlier work this paper cites.
B. Mao, S. Wu, and L. Duan, “Improving the ssd performance by exploiting request characteristics and internal parallelism,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 37, no. 2, pp. 472–484, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Hadian and T. Heinis, “Towards batch-processing on cold storage devices,” in 2018 IEEE 34th International Conference on Data Engineering Workshops (ICDEW) . IEEE, 2018, pp. 134–139
2018
Earlier work this paper cites.
A. Hadian and T. Heinis, “Towards batch-processing on cold storage devices,” in 2018 IEEE 34th International Conference on Data Engineering Workshops (ICDEW) . IEEE, 2018, pp. 134–139
2018
Earlier work this paper cites.
S. Liang, Y. Wang, Y. Lu, Z. Yang, H. Li, and X. Li, “Cognitive { \{ SSD } \} : A deep learning engine for { \{ In-Storage } \} data retrieval,” in 2019 USENIX Annual Technical Conference (USENIX ATC 19) , 2019, pp. 395–410
2019
Earlier work this paper cites.
J. Zhang, M. Kwon, H. Kim, H. Kim, and M. Jung, “Flashgpu: Placing new flash next to gpu cores,” in Proceedings of the 56th Annual Design Automation Conference 2019 , 2019, pp. 1–6
2019
Earlier work this paper cites.
S. Liang, Y. Wang, Y. Lu, Z. Yang, H. Li, and X. Li, “Cognitive { \{ SSD } \} : A deep learning engine for { \{ In-Storage } \} data retrieval,” in 2019 USENIX Annual Technical Conference (USENIX ATC 19) , 2019, pp. 395–410
2019
Earlier work this paper cites.
J. Zhang, M. Kwon, H. Kim, H. Kim, and M. Jung, “Flashgpu: Placing new flash next to gpu cores,” in Proceedings of the 56th Annual Design Automation Conference 2019 , 2019, pp. 1–6
2019
Earlier work this paper cites.
J. Kwak, S. Lee, K. Park, J. Jeong, and Y. H. Song, “Cosmos+ openssd: Rapid prototype for flash storage systems,” ACM Transactions on Storage (TOS) , vol. 16, no. 3, pp. 1–35, 2020
2020
Earlier work this paper cites.
K. Myung, S. Kim, H. Y. Yeom, and J. Park, “Efficient and scalable external sort framework for nvme ssd,” IEEE Transactions on Computers , vol. 70, no. 12, pp. 2211–2217, 2020
2020
Earlier work this paper cites.
J. Zhang and M. Jung, “Zng: Architecting gpu multi-processors with new flash for scalable data analysis,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2020, pp. 1064–1075
2020
Earlier work this paper cites.
J. Kwak, S. Lee, K. Park, J. Jeong, and Y. H. Song, “Cosmos+ openssd: Rapid prototype for flash storage systems,” ACM Transactions on Storage (TOS) , vol. 16, no. 3, pp. 1–35, 2020
2020
Earlier work this paper cites.
K. Myung, S. Kim, H. Y. Yeom, and J. Park, “Efficient and scalable external sort framework for nvme ssd,” IEEE Transactions on Computers , vol. 70, no. 12, pp. 2211–2217, 2020
2020
Earlier work this paper cites.
J. Zhang and M. Jung, “Zng: Architecting gpu multi-processors with new flash for scalable data analysis,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2020, pp. 1064–1075
2020
Earlier work this paper cites.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 12, no. 5, pp. 1–32, 2021
2021
Earlier work this paper cites.
B. Chen, T. Dao, E. Winsor, Z. Song, A. Rudra, and C. Ré, “Scatterbrain: Unifying sparse and low-rank attention,” Advances in Neural Information Processing Systems , vol. 34, pp. 17 413–17 426, 2021
2021
Earlier work this paper cites.
S. Chen, S. Huang, S. Pandey, B. Li, G. R. Gao, L. Zheng, C. Ding, and H. Liu, “Et: re-thinking self-attention for transformer models on gpus,” in Proceedings of the international conference for high performance computing, networking, storage and analysis , 2021, pp. 1–18
2021
Earlier work this paper cites.
J. Markussen, L. B. Kristiansen, P. Halvorsen, H. Kielland-Gyrud, H. K. Stensland, and C. Griwodz, “Smartio: Zero-overhead device sharing through pcie networking,” ACM Transactions on Computer Systems , vol. 38, no. 1–2, jul 2021
2021
Earlier work this paper cites.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 12, no. 5, pp. 1–32, 2021
2021
Earlier work this paper cites.
B. Chen, T. Dao, E. Winsor, Z. Song, A. Rudra, and C. Ré, “Scatterbrain: Unifying sparse and low-rank attention,” Advances in Neural Information Processing Systems , vol. 34, pp. 17 413–17 426, 2021
2021
Earlier work this paper cites.
S. Chen, S. Huang, S. Pandey, B. Li, G. R. Gao, L. Zheng, C. Ding, and H. Liu, “Et: re-thinking self-attention for transformer models on gpus,” in Proceedings of the international conference for high performance computing, networking, storage and analysis , 2021, pp. 1–18
2021
Earlier work this paper cites.
J. Markussen, L. B. Kristiansen, P. Halvorsen, H. Kielland-Gyrud, H. K. Stensland, and C. Griwodz, “Smartio: Zero-overhead device sharing through pcie networking,” ACM Transactions on Computer Systems , vol. 38, no. 1–2, jul 2021
2021
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Earlier work this paper cites.
Y. Lee, J. Chung, and M. Rhu, “Smartsage: training large-scale graph neural networks using in-storage processing architectures,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , 2022, pp. 932–945
2022
Earlier work this paper cites.
N. Mansouri Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. Senol Cali, C. Firtina, H. Mao, N. Almadhoun Alserr et al. , “Genstore: A high-performance in-storage processing system for genome sequence analysis,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2022, pp. 635–654
2022
Earlier work this paper cites.
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,” in International Conference on Learning Representations , 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:252668917
2022
Earlier work this paper cites.
M. Soltaniyeh, V. Lagrange Moutinho Dos Reis, M. Bryson, X. Yao, R. P. Martin, and S. Nagarakatte, “Near-storage processing for solid state drive based recommendation inference with smartssds®,” in Proceedings of the 2022 ACM/SPEC on International Conference on Performance Engineering , 2022, pp. 177–186
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory-based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2022, pp. 1071–1085
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Cited alongside, same era.
Y. Lee, J. Chung, and M. Rhu, “Smartsage: training large-scale graph neural networks using in-storage processing architectures,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , 2022, pp. 932–945
2022
Cited alongside, same era.
N. Mansouri Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. Senol Cali, C. Firtina, H. Mao, N. Almadhoun Alserr et al. , “Genstore: A high-performance in-storage processing system for genome sequence analysis,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2022, pp. 635–654
2022
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava, “Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Pan, Y. An, S. Liang, B. Mao, M. Zhang, Q. Li, M. Jung, and J. Zhang, “Flagger: Cooperative acceleration for large-scale cross-silo federated learning aggregation,” in Proceedings of the 51th Annual International Symposium on Computer Architecture , 2024, pp. 915–930
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,” in International Conference on Learning Representations , 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:252668917
2022
Cited alongside, same era.
M. Soltaniyeh, V. Lagrange Moutinho Dos Reis, M. Bryson, X. Yao, R. P. Martin, and S. Nagarakatte, “Near-storage processing for solid state drive based recommendation inference with smartssds®,” in Proceedings of the 2022 ACM/SPEC on International Conference on Performance Engineering , 2022, pp. 177–186
2022
Cited alongside, same era.
2022
Cited alongside, same era.
M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory-based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2022, pp. 1071–1085
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Choi, J. Park, K. Kyung, N. S. Kim, and J. H. Ahn, “Unleashing the potential of pim: Accelerating large batched inference of transformer-based generative models,” IEEE Computer Architecture Letters , 2023
2023
Cited alongside, same era.
G. Haas and V. Leis, “What modern nvme storage can do, and how to exploit it: high-performance i/o for high-performance storage engines,” Proceedings of the VLDB Endowment , vol. 16, no. 9, pp. 2090–2102, 2023
2023
Cited alongside, same era.
S.-H. Kim, J. Shim, E. Lee, S. Jeong, I. Kang, and J.-S. Kim, “ { \{ NVMeVirt } \} : A versatile software-defined virtual { \{ NVMe } \} device,” in 21st USENIX Conference on File and Storage Technologies (FAST 23) , 2023, pp. 379–394
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Wang, X. Pan, Y. An, J. Zhang, and G. Reinman, “Beacongnn: Large-scale gnn acceleration with out-of-order streaming in-storage computing,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2024, pp. 330–344
2024
Closest in time.
Y. Wu, Z. Wang, and W. D. Lu, “Pim gpt a hybrid process in memory accelerator for autoregressive transformers,” npj Unconventional Computing , vol. 1, no. 1, p. 4, 2024
2024
Closest in time.
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, S. Zhong, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” ACM Trans. Knowl. Discov. Data , vol. 18, no. 6, apr 2024. [Online]. Available: https://doi.org/10.1145/3649506
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, S. Liu, R. Chen, B. Kailkhura, B. Chen, and A. Wang, “Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache,” Proceedings of Machine Learning and Systems , vol. 6, pp. 381–394, 2024
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “ { \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 193–210
2024
Closest in time.
N. Express, “Nvm express base specification 2.0d.” [Online]. Available: https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0d-2024.01.11-Ratified.pdf
2024
Closest in time.
Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, and L. Mai, “ { \{ ServerlessLLM } \} : { \{ Low-Latency } \} serverless inference for large language models,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 135–153
2024
Closest in time.
B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, J. Deng, X. Yang, Z. Yu, and P. Zuo, “ { \{ Cost-Efficient } \} large language model serving for multi-turn conversations with { \{ CachedAttention } \} ,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) , 2024, pp. 111–126
2024
Closest in time.
2024
Closest in time.
G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 722–737
2024
Closest in time.
2024
Closest in time.
K. Hong, G. Dai, J. Xu, Q. Mao, X. Li, J. Liu, Y. Dong, Y. Wang et al. , “Flashdecoding++: Faster large language model inference with asynchronization, flat gemm optimization, and heuristics,” Proceedings of Machine Learning and Systems , vol. 6, pp. 148–161, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Lee, J. Lee, J. Seo, and J. Sim, “InfiniGen: Efficient generative inference of large language models with dynamic KV cache management,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 155–172. [Online]. Available: https://www.usenix.org/conference/osdi24/presentation/lee
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava, “Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Pan, Y. An, S. Liang, B. Mao, M. Zhang, Q. Li, M. Jung, and J. Zhang, “Flagger: Cooperative acceleration for large-scale cross-silo federated learning aggregation,” in Proceedings of the 51th Annual International Symposium on Computer Architecture , 2024, pp. 915–930
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Wang, X. Pan, Y. An, J. Zhang, and G. Reinman, “Beacongnn: Large-scale gnn acceleration with out-of-order streaming in-storage computing,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2024, pp. 330–344
2024
Closest in time.
Y. Wu, Z. Wang, and W. D. Lu, “Pim gpt a hybrid process in memory accelerator for autoregressive transformers,” npj Unconventional Computing , vol. 1, no. 1, p. 4, 2024
2024
Closest in time.
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, S. Zhong, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” ACM Trans. Knowl. Discov. Data , vol. 18, no. 6, apr 2024. [Online]. Available: https://doi.org/10.1145/3649506
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, S. Liu, R. Chen, B. Kailkhura, B. Chen, and A. Wang, “Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache,” Proceedings of Machine Learning and Systems , vol. 6, pp. 381–394, 2024
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “ { \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 193–210
2024
Closest in time.