Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate outstanding performance in various tasks in machine learning and have thus become one of the most important workloads in today's computing landscape.
M. Marcus, G. Kim, M. A. Marcinkiewicz, R. MacIntyre, A. Bies, M. Ferguson, K. Katz, and B. Schasberger, “The penn treebank: Annotating predicate argument structure,” in Proceedings of the Workshop on Human Language Technology , 1994
1994
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “Dadiannao: A machine-learning supercomputer,” in 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2014
2014
Earlier work this paper cites.
J. S. S. T. Association, JEDEC Standard JESD235A: High Bandwidth Memory (HBM) DRAM , JEDEC, Virginia, USA, 2015
2015
Earlier work this paper cites.
Y. Kim, W. Yang, and O. Mutlu, “Ramulator: A fast and extensible dram simulator,” IEEE Computer architecture letters (CAL) , 2015
2015
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
M. Alwani, H. Chen, M. Ferdman, and P. Milder, “Fused-layer cnn accelerators,” in 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2016
2016
Earlier work this paper cites.
Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” in ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G.-Y. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-x: An accelerator for sparse neural networks,” in 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2016
2016
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” in ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) , 2017
2017
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
E. Nurvitadhi, G. Venkatesh, J. Sim, D. Marr, R. Huang, J. Ong Gee Hock, Y. T. Liew, K. Srivatsan, D. Moss, S. Subhaschandra, and G. Boudoukh, “Can fpgas beat gpus in accelerating next-generation deep neural networks?” in ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA) , 2017
2017
Earlier work this paper cites.
M. O’Connor, N. Chatterjee, D. Lee, J. Wilson, A. Agrawal, S. W. Keckler, and W. J. Dally, “Fine-grained dram: Energy-efficient dram for extreme bandwidth systems,” in 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2017
2017
Earlier work this paper cites.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
C. Eckert, X. Wang, J. Wang, A. Subramaniyan, R. Iyer, D. Sylvester, D. Blaauw, and R. Das, “Neural cache: Bit-serial in-cache acceleration of deep neural networks,” in ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) , 2018
2018
Earlier work this paper cites.
J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. K. Reinhardt, A. M. Caulfield, E. S. Chung, and D. Burger, “A configurable cloud-scale dnn processor for real-time ai,” in ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) , 2018
2018
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Earlier work this paper cites.
E. Park, D. Kim, and S. Yoo, “Energy-efficient neural network accelerator based on outlier-aware low-precision computation,” in ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) , 2018
2018
Earlier work this paper cites.
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, J. K. Kim, V. Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,” in ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , 2019
2019
Earlier work this paper cites.
S. Jain, S. Venkataramani, V. Srinivasan, J. Choi, K. Gopalakrishnan, and L. Chang, “Biscaled-dnn: Quantizing long-tailed datastructures with two scale factors for deep neural networks,” in 56th ACM/IEEE Design Automation Conference (DAC) , 2019
2019
Earlier work this paper cites.
Y. Kwon, Y. Lee, and M. Rhu, “Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,” in 52nd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2019
2019
Cited alongside, same era.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in International Conference on Learning Representations (ICLR) , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
B. Darvish Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers, K. Ovtcharov, A. Vinogradsky, S. Massengill, L. Yang, R. Bittner, A. Forin, H. Zhu, T. Na, P. Patel, S. Che, L. Chand Koppaka, X. SONG, S. Som, K. Das, S. T, S. Reinhardt, S. Lanka, E. Chung, and D. Burger, “Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
X. Wei, Y. Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu, “Outlier suppression: Pushing the limit of low-bit transformer language models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
A. Yazdanbakhsh, A. Moradifirouzabadi, Z. Li, and M. Kang, “Sparse attention acceleration with synergistic in-memory pruning and on-chip recomputation,” in 55th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2022
2022
Later among the works it cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI) , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Intel. (2020) Intel neural compressor. [Online]. Available: https://intel.github.io/neural-compressor
2020
Cited alongside, same era.
L. Ke, U. Gupta, B. Y. Cho, D. Brooks, V. Chandra, U. Diril, A. Firoozshahian, K. Hazelwood, B. Jia, H.-H. S. Lee, M. Li, B. Maher, D. Mudigere, M. Naumov, M. Schatz, M. Smelyanskiy, X. Wang, B. Reagen, C.-J. Wu, M. Hempstead, and X. Zhang, “Recnmp: Accelerating personalized recommendation with near-memory processing,” in ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , 2020
2020
Cited alongside, same era.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020
2020
Cited alongside, same era.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-BERT: hessian based ultra low precision quantization of BERT,” in The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI) , 2020
2020
Cited alongside, same era.
Z. Song, B. Fu, F. Wu, Z. Jiang, L. Jiang, N. Jing, and X. Liang, “Drq: Dynamic region-based quantization for deep neural network acceleration,” in ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , 2020
2020
Cited alongside, same era.
T. Tambe, E.-Y. Yang, Z. Wan, Y. Deng, V. Janapa Reddi, A. Rush, D. Brooks, and G.-Y. Wei, “Algorithm-hardware co-design of adaptive floating-point encodings for resilient deep learning inference,” in 57th ACM/IEEE Design Automation Conference (DAC) , 2020
2020
Cited alongside, same era.
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos, “Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference,” in 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
D. Chen, H. He, H. Jin, L. Zheng, Y. Huang, X. Shen, and X. Liao, “Metanmp: Leveraging cartesian-like product to accelerate hgnns with near-memory processing,” in ACM/IEEE 50th Annual International Symposium on Computer Architecture (ISCA) , 2023
2023
Later among the works it cites.
B. Darvish Rouhani, R. Zhao, V. Elango, R. Shafipour, M. Hall, M. Mesmakhosroshahi, A. More, L. Melnick, M. Golub, G. Varatkar, L. Shao, G. Kolhe, D. Melts, J. Klar, R. L’Heureux, M. Perry, D. Burger, E. Chung, Z. S. Deng, S. Naghshineh, J. Park, and M. Naumov, “With shared microexponents, a little shifting goes a long way,” in ACM/IEEE 50th Annual International Symposium on Computer Architecture (ISCA) , 2023
2023
Later among the works it cites.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Later among the works it cites.
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “OPTQ: Accurate quantization for generative pre-trained transformers,” in International Conference on Learning Representations (ICLR) , 2023
2023
Later among the works it cites.
Google. (2023) Gemini. [Online]. Available: https://gemini.google.com
2023
Later among the works it cites.
C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, Y. Liu, M. Guo, and Y. Zhu, “Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization,” in ACM/IEEE 50th Annual International Symposium on Computer Architecture (ISCA) , 2023
2023
Later among the works it cites.
N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. A. Patterson, “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in ACM/IEEE 50th Annual International Symposium on Computer Architecture (ISCA) , 2023
2023
Later among the works it cites.
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han, “Awq: Activation-aware weight quantization for llm compression and acceleration,” arXiv , 2023
2023
Later among the works it cites.
H. Liu, L. Zheng, Y. Huang, C. Liu, X. Ye, J. Yuan, X. Liao, H. Jin, and J. Xue, “Accelerating personalized recommendation with cross-level near-memory processing,” in ACM/IEEE 50th Annual International Symposium on Computer Architecture (ISCA) , 2023
2023
Later among the works it cites.
OCP Microscaling Formats (MX) Specification Version 1.0 , Open Compute Project, 9 2023
2023
Later among the works it cites.
OpenAI. (2023) Chatgpt. [Online]. Available: https://openai.com/chatgpt
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning (ICML) , 2023
2023
Later among the works it cites.
Synopsys. (2023) Design compiler - synopsys. [Online]. Available: https://www.synopsys.com/implementation-and-signoff/rtl-synthesis-test/dc-ultra.html
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning (ICML) , 2023
2023
Later among the works it cites.