Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have emerged due to their capability to generate high-quality content across diverse contexts.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learners,” in NeurIPS , 2020, pp. 1877–1901, doi:10.5555/3495724.3495883
1901
Earlier work this paper cites.
P. Grun, “Introduction to infiniband for end users,” White paper, InfiniBand Trade Association , vol. 55, 2010
2010
Earlier work this paper cites.
A. Shafaei, Y. Wang, X. Lin, and M. Pedram, “FinCACTI: Architectural Analysis and Modeling of Caches with Deeply-Scaled FinFET Devices,” in 2014 IEEE Computer Society Annual Symposium on VLSI , 2014, pp. 290–295, doi:10.1109/ISVLSI.2014.94
2014
Earlier work this paper cites.
D. Zhang, N. Jayasena, A. Lyashevsky, J. L. Greathouse, L. Xu, and M. Ignatowski, “TOP-PIM: Throughput-oriented programmable processing in memory,” in Proceedings of the 23rd international symposium on High-performance parallel and distributed computing , 2014, pp. 85–98, doi:10.1145/2600212.2600213
2014
Earlier work this paper cites.
L. T. Clark, V. Vashishtha, L. Shifren, A. Gujja, S. Sinha, B. Cline, C. Ramamurthy, and G. Yeric, “ASAP7: A 7-nm FinFET Predictive Process Design Kit,” Microelectronics Journal , vol. 53, pp. 105–115, 2016, doi:10.1016/j.mejo.2016.04.006
2016
Earlier work this paper cites.
S. Wu, C. Lin, M. Chiang, J. Liaw, J. Cheng, S. Yang, C. Tsai, P. Chen, T. Miyashita, C. Chang, V. Chang, K. Pan, J. Chen, Y. Mor, K. Lai, C. Liang, H. Chen, S. Chang, C. Lin, C. Hsieh, R. Tsui, C. Yao, C. Chen, R. Chen, C. Lee, H. Lin, C. Chang, K. Chen, M. Tsai, K. Chen, Y. Ku, and S. Jang, “A 7nm CMOS Platform Technology Featuring 4th Generation FinFET Transistors with a 0.027um2 High Density 6-T SRAM cell for Mobile SoC Applications,” in IEEE International Electron Devices Meeting , 2016, pp. 2.6.1–2.6.4, doi:10.1109/IEDM.2016.7838333
2016
Earlier work this paper cites.
J. Chang, Y. Chen, W. Chan, S. P. Singh, H. Cheng, H. Fujiwara, J. Lin, K. Lin, J. Hung, R. Lee, H. Liao, J. Liaw, Q. Li, C. Lin, M. Chiang, and S. Wu, “A 7nm 256Mb SRAM in High-K Metal-Gate FinFET Technology with Write-Assist Circuitry for Low-VMIN Applications,” in IEEE International Solid-State Circuits Conference , 2017, pp. 206–207, doi:10.1109/ISSCC.2017.7870333
2017
Earlier work this paper cites.
H. Jun, J. Cho, K. Lee, H.-Y. Son, K. Kim, H. Jin, and K. Kim, “HBM (High Bandwidth Memory) DRAM Technology and Architecture,” in 2017 IEEE International Memory Workshop (IMW) , 2017, pp. 1–4, doi:10.1109/IMW.2017.7939084
2017
Earlier work this paper cites.
M. O’Connor, N. Chatterjee, D. Lee, J. Wilson, A. Agrawal, S. W. Keckler, and W. J. Dally, “Fine-Grained DRAM: Energy-Efficient DRAM for Extreme Bandwidth Systems,” in MICRO , 2017, p. 41–54, doi:10.1145/3123939.3124545
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in NeurIPS , 2017. [Online]. Available: https://dl.acm.org/doi/10.5555/3295222.3295349
2017
Earlier work this paper cites.
IEEE, “International Roadmap for Devices and Systems: 2018,” Tech. Rep., 2018. [Online]. Available: https://irds.ieee.org/editions/2018/
2018
Earlier work this paper cites.
W. Jeong, S. Maeda, H. Lee, K. Lee, T. Lee, D. Park, B. Kim, J. Do, T. Fukai, D. Kwon, K. Nam, W. Rim, M. Jang, H. Kim, Y. Lee, J. Park, E. Lee, D. Ha, C. Park, H. Cho, S. Jung, and H. Kang, “True 7nm Platform Technology featuring Smallest FinFET and Smallest SRAM cell by EUV, Special Constructs and 3rd Generation Single Diffusion Break,” in IEEE Symposium on VLSI Technology , 2018, pp. 59–60, doi:10.1109/VLSIT.2018.8510682
2018
Earlier work this paper cites.
S. Lee, K. Lee, M. Sung, M. Alian, C. Kim, W. Cho, R. Oh, S. O, J. Ahn, and N. S. Kim, “3D-Xpath: High-density Managed DRAM Architecture with Cost-effective Alternative Paths for Memory Transactions,” in PACT , 2018, doi:10.1145/3243176.3243191
2018
Earlier work this paper cites.
T. Song, J. Jung, W. Rim, H. Kim, Y. Kim, C. Park, J. Do, S. Park, S. Cho, H. Jung, B. Kwon, H. Choi, J. Choi, and J. S. Yoon, “A 7nm FinFET SRAM Using EUV Lithography with Dual Write-Driver-Assist Circuitry for Low-Voltage Applications,” in IEEE International Solid-State Circuits Conference , 2018, doi:10.1109/ISSCC.2018.8310252
2018
Earlier work this paper cites.
F. Devaux, “The True Processing In Memory Accelerator,” in 2019 IEEE Hot Chips 31 Symposium (HCS) , 2019, doi:10.1109/HOTCHIPS.2019.8875680
2019
Earlier work this paper cites.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J.-H. Park, S. Lee, K. Park, J. W. Lee, and D.-K. Jeong, “ A 3 A^{3} : Accelerating Attention Mechanisms in Neural Networks with Approximation,” in HPCA , 2020, pp. 328–341, doi:10.1109/HPCA47549.2020.00035
2020
Earlier work this paper cites.
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos, “GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference,” in MICRO , 2020, pp. 811–824, doi:10.1109/MICRO50266.2020.00071
2020
Earlier work this paper cites.
T. J. Ham, Y. Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks,” in ISCA , 2021, pp. 692–705, doi:10.1109/ISCA52012.2021.00060
2021
Earlier work this paper cites.
N. P. Jouppi, D. H. Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. C. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. A. Patterson, “Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product,” in ISCA , 2021, pp. 1–14, doi:10.1109/ISCA52012.2021.00010
2021
Earlier work this paper cites.
S. Lee, S.-h. Kang, J. Lee, H. Kim, E. Lee, S. Seo, H. Yoon, S. Lee, K. Lim, H. Shin, J. Kim, O. Seongil, A. Iyer, D. Wang, K. Sohn, and N. S. Kim, “Hardware Architecture and Software Stack for PIM based on Commercial DRAM Technology: Industrial Product,” in ISCA , 2021, doi:10.1109/ISCA52012.2021.00013
2021
Earlier work this paper cites.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture,” in MICRO , 2021, p. 977–991, doi:10.1145/3466752.3480125
2021
Cited alongside, same era.
N. Park, S. Ryu, J. Kung, and J.-J. Kim, “High-throughput Near-memory Processing on CNNs with 3D HBM-like Memory,” ACM Transactions on Design Automation of Electronic Systems , vol. 26, no. 6, pp. 1–20, 2021, doi:10.1145/3460971
2021
Cited alongside, same era.
H. Wang, Z. Zhang, and S. Han, “SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning,” in HPCA , 2021, pp. 97–110, doi:10.1109/HPCA51647.2021.00018
2021
Cited alongside, same era.
M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V. Lin, J. Du, S. Iyer, R. Pasunuru, G. Anantharaman, X. Li, S. Chen, H. Akin, M. Baines, L. Martin, X. Zhou, P. S. Koura, B. O’Horo, J. Wang, L. Zettlemoyer, M. Diab, Z. Kozareva, and V. Stoyanov, “Efficient Large Scale Language Modeling with Mixtures of Experts,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 11 699–11 732, doi:10.18653/v1/2022.emnlp-main.804
M. Zhou, W. Xu, J. Kang, and T. Rosing, “TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for Transformer,” in HPCA , 2022, doi:10.1109/HPCA53966.2022.00082
2022
Later among the works it cites.
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Singapore, 2023, pp. 4895–4901, doi:10.18653/v1/2023.emnlp-main.298
2023
Later among the works it cites.
J. Choi, J. Park, K. Kyung, N. S. Kim, and J. Ahn, “Unleashing the Potential of PIM: Accelerating Large Batched Inference of Transformer-Based Generative Models,” IEEE Computer Architecture Letters , vol. 22, pp. 113–116, 2023, doi:10.1109/LCA.2023.3305386
2023
Later among the works it cites.
S. R. Group, “Ramulator 2.0 — GitHub Repository,” 2023. [Online]. Available: https://github.com/CMU-SAFARI/ramulator2
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. P. Bosma, Z. Zhou, T. Wang, E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. Le, Y. Wu, Z. Chen, and C. Cui, “GLaM: Efficient Scaling of Language Models with Mixture-of-Experts,” in Proceedings of the 39th International Conference on Machine Learning , vol. 162, 2022, pp. 5547–5569. [Online]. Available: https://proceedings.mlr.press/v162/du22c.html
2022
Cited alongside, same era.
H. Fan, T. Chau, S. I. Venieris, R. Lee, A. Kouris, W. Luk, N. D. Lane, and M. S. Abdelfattah, “Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design,” in MICRO , 2022, pp. 599–615, doi:10.1109/MICRO56248.2022.00050
2022
Cited alongside, same era.
C. Fang, A. Zhou, and Z. Wang, “An Algorithm–Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 30, no. 11, pp. 1573–1586, 2022, doi:10.1109/TVLSI.2022.3197282
2022
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022. [Online]. Available: http://jmlr.org/papers/v23/21-0998.html
2022
Cited alongside, same era.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation,” in MICRO , 2022, p. 616–630, doi:10.1109/MICRO56248.2022.00051
2022
Cited alongside, same era.
JEDEC, “High Bandwidth Memory DRAM (HBM3),” 2022
2022
Cited alongside, same era.
S. Lee, K. Kim, S. Oh, J. Park, G. Hong, D. Ka, K. Hwang, J. Park, K. Kang, J. Kim, J. Jeon, N. Kim, Y. Kwon, K. Vladimir, W. Shin, J. Won, M. Lee, H. Joo, H. Choi, J. Lee, D. Ko, Y. Jun, K. Cho, I. Kim, C. Song, C. Jeong, D. Kwon, J. Jang, I. Park, J. Chun, and J. Cho, “A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6-based Accelerator-in-Memory supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep-Learning Applications,” in 2022 IEEE International Solid-State Circuits Conference , vol. 65, 2022, pp. 1–3, doi:10.1109/ISSCC42614.2022.9731711
2022
Cited alongside, same era.
Z. Li, S. Ghodrati, A. Yazdanbakhsh, H. Esmaeilzadeh, and M. Kang, “Accelerating Attention through Gradient-Based Learned Runtime Pruning,” in ISCA , 2022, p. 902–915, doi:10.1145/3470496.3527423
2022
Cited alongside, same era.
Later among the works it cites.
C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, Y. Liu, M. Guo, and Y. Zhu, “OliVe: Accelerating Large Language Models via Hardware-Friendly Outlier-Victim Pair Quantization,” in ISCA , 2023, doi:10.1145/3579371.3589038
2023
Later among the works it cites.
H. Huang, N. Ardalani, A. Sun, L. Ke, H.-H. S. Lee, A. Sridhar, S. Bhosale, C.-J. Wu, and B. Lee, “Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference,” 2023,” doi:10.48550/arXiv.2303.06182
2023
Later among the works it cites.
S.-C. Kao, S. Subramanian, G. Agrawal, A. Yazdanbakhsh, and T. Krishna, “FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks,” in ASPLOS, Volume 2 , 2023, p. 295–310, doi:10.1145/3575693.3575747
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, p. 611–626, doi:10.1145/3600006.3613165
2023
Later among the works it cites.
H. Luo, Y. C. Tuğrul, F. N. Bostancı, A. Olgun, A. G. Yağlıkçı, and O. Mutlu, “Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator,” IEEE Computer Architecture Letters , vol. 23, no. 1, 2023, doi:10.1109/LCA.2023.3333759
2023
Later among the works it cites.
OpenAI, “GPT-4 Technical Report,” 2023
2023
Later among the works it cites.
Y. Qin, Y. Wang, D. Deng, Z. Zhao, X. Yang, L. Liu, S. Wei, Y. Hu, and S. Yin, “FACT: FFN-Attention Co-Optimized Transformer Architecture with Eager Correlation Prediction,” in ISCA , 2023, doi:10.1145/3579371.3589057
2023
Later among the works it cites.
SK hynix, “Advanced Packaging Technology for Beyond Memory,” p. 29, 2023. [Online]. Available: https://www.theise.org/wp-content/uploads/2023/10/Tutorial1-4_%EC%86%90%ED%98%B8%EC%98%81%EC%88%98%EC%84%9D%EB%8B%98_SK%ED%95%98%EC%9D%B4%EB%8B%89%EC%8A%A4.pdf
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom, “Llama 2: Open Foundation and Fine-Tuned Chat Models,” 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2307.09288
2023
Later among the works it cites.
G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing,” in ASPLOS, Volume 3 , 2024, p. 722–737, doi:10.1145/3620666.3651380
2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mixtral of Experts,” 2024,” doi:10.48550/arXiv.2401.04088
2024
Closest in time.
Meta, “Llama3,” 2024. [Online]. Available: https://llama.meta.com/llama3/
2024
Closest in time.
J. Park, J. Choi, K. Kyung, M. J. Kim, Y. Kwon, N. S. Kim, and J. Ahn, “AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference,” in ASPLOS, Volume 2 , 2024, p. 103–119, doi:10.1145/3620665.3640422
2024
Closest in time.
P. Patel, E. Choukse, C. Zhang, A. Shah, Í. Goiri, S. Maleki, and R. Bianchini, “Splitwise: Efficient Generative LLM Inference Using Phase Splitting,” in ISCA , 2024, pp. 118–132, doi:10.1109/ISCA59077.2024.00019
2024
Closest in time.
S. Yun, H. Nam, K. Kyung, J. Park, B. Kim, Y. Kwon, E. Lee, and J. Ahn, “CLAY: CXL-based Scalable NDP Architecture Accelerating Embedding Layers,” in Proceedings of the 38th ACM International Conference on Supercomputing , 2024, p. 338–351, doi:10.1145/3650200.3656595
2024
Closest in time.