Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown remarkable performance across a wide range of applications, often outperforming human experts.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
N. P. Jouppi, , C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P. l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA) , 2017
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Parashar, P. Raina, Y. S. Shao, Y.-H. Chen, V. A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A systematic approach to dnn accelerator evaluation,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2019, pp. 304–315
2019
Earlier work this paper cites.
A. Parashar, P. Raina, Y. S. Shao, Y.-H. Chen, V. A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A systematic approach to dnn accelerator evaluation,” in 2019 IEEE international symposium on performance analysis of systems and software (ISPASS) . IEEE, 2019, pp. 304–315
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” 2019
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
Y. N. Wu, J. S. Emer, and V. Sze, “Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs,” in IEEE/ACM International Conference On Computer Aided Design (ICCAD) , 2019
2019
Earlier work this paper cites.
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. R. Tallent, and K. J. Barker, “Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect,” IEEE Transactions on Parallel and Distributed Systems , vol. 31, no. 1, p. 94–110, Jan. 2020. [Online]. Available: http://dx.doi.org/10.1109/TPDS.2019.2928289
2019
Earlier work this paper cites.
M. Brysbaert, “How many words do we read per minute? a review and meta-analysis of reading rate,” Journal of Memory and Language , vol. 109, p. 104047, 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0749596X19300786
2019
Earlier work this paper cites.
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Rashidi, S. Sridharan, S. Srinivasan, and T. Krishna, “ASTRA-SIM: Enabling sw/hw co-design exploration for distributed dl training platforms,” in IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2020
2020
Earlier work this paper cites.
A. Samajdar, J. M. Joseph, Y. Zhu, P. Whatmough, M. Mattina, and T. Krishna, “A systematic methodology for characterizing scalability of dnn accelerators using scale-sim,” in 2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2020, pp. 58–68
2020
Earlier work this paper cites.
H. Kwon, P. Chatarasi, V. Sarkar, T. Krishna, M. Pellauer, and A. Parashar, “Maestro: A data-centric approach to understand reuse, performance, and hardware cost of dnn mappings,” IEEE Micro , vol. 40, no. 3, pp. 20–29, 2020
2020
Earlier work this paper cites.
K. Rocki, D. Van Essendelft, I. Sharapov, R. Schreiber, M. Morrison, V. Kibardin, A. Portnoy, J. F. Dietiker, M. Syamlal, and M. James, “Fast stencil-code computation on a wafer-scale processor,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–14
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
SambaNova, “Sambanova whitepaper,” 2021. [Online]. Available: https://sambanova.ai/wp-content/uploads/2021/06/SambaNova%5FRDA%5FWhitepaper%5FEnglish.pdf
2021
Earlier work this paper cites.
V. Kandiah, S. Peverelle, M. Khairy, J. Pan, A. Manjunath, T. G. Rogers, T. M. Aamodt, and N. Hardavellas, “Accelwattch: A power modeling framework for modern gpus,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 738–753. [Online]. Available: https://doi.org/10.1145/3466752.3480063
2021
Earlier work this paper cites.
M. Khani, M. Ghobadi, M. Alizadeh, Z. Zhu, M. Glick, K. Bergman, A. Vahdat, B. Klenk, and E. Ebrahimi, “Sip-ml: high-bandwidth optical network interconnects for machine learning training,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference , ser. SIGCOMM ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 657–675. [Online]. Available: https://doi.org/10.1145/3452296.3472900
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,” 2022
2022
Earlier work this paper cites.
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A survey of quantization methods for efficient neural network inference,” in Low-Power Computer Vision . Chapman and Hall/CRC, 2022, pp. 291–326
2022
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” 2022
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Earlier work this paper cites.
C. W. Kim, “Guiding text generation with constrained beam search in transformers,” Mar 2022. [Online]. Available: https://huggingface.co/blog/constrained-beam-search
2022
Earlier work this paper cites.
S. Li, F. Xue, C. Baranwal, Y. Li, and Y. You, “Sequence parallelism: Long sequence training from system perspective,” 2022
2022
Earlier work this paper cites.
“Groqchip.” [Online]. Available: https://groq.com/wp-content/uploads/2022/10/GroqChip\texttrademark-Processor-Product-Brief-v1.5.pdf
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Google, “Introducing gemini: Google’s most capable ai model yet,” 2023. [Online]. Available: https://blog.google/technology/ai/google-gemini-ai/
2023
Earlier work this paper cites.
E. Caballero, K. Gupta, I. Rish, and D. Krueger, “Broken neural scaling laws,” 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” PMLR, pp. 38 087–38 099, 2023
2023
Cited alongside, same era.
E. Frantar and D. Alistarh, “Sparsegpt: Massive language models can be accurately pruned in one-shot,” in International Conference on Machine Learning . PMLR, 2023, pp. 10 323–10 337
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. R. Bambhaniya, A. Yazdanbakhsh, S. Subramanian, S.-C. Kao, S. Agrawal, U. Evci, and T. Krishna, “Progressive gradient flow for robust n:m sparsity training in transformers,” 2024
2024
Closest in time.
G. Jeong, P.-A. Tsai, A. R. Bambhaniya, S. W. Keckler, and T. Krishna, “Abstracting sparse dnn acceleration via structured sparse tensor decomposition,” 2024
2024
Closest in time.
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer, “Speculative decoding with big little decoder,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
NVIDIA, “Nvidia. tensorrt-llm.” 2023. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Sun, X. Liu, Y. Gong, Y. Zhang, D. Jiang, L. Yang, and N. Duan, “Allies: Prompting large language model with beam search,” in Findings of EMNLP , 2023
2023
Cited alongside, same era.
M. Agarwal, A. Qureshi, N. Sardana, L. Li, J. Quevedo, and D. Khudia, “Llm inference performance engineering: Best practices,” Oct 2023. [Online]. Available: https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices
2023
Cited alongside, same era.
H. Wu, Y. Gan, F. Yuan, J. Ma, W. Zhu, Y. Xu, H. Zhu, Y. Zhu, X. Liu, and J. Gu, “Efficient llm inference solution on intel gpu,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
“Streamlining AI Inference Performance and Deployment with NVIDIA TensorRT-LLM Chunked Prefill — NVIDIA Technical Blog — developer.nvidia.com,” https://developer.nvidia.com/blog/streamlining-ai-inference-performance-and-deployment-with-nvidia-tensorrt-llm-chunked-prefill/ , 2024, [Accessed 22-11-2024]
2024
Closest in time.
“Demystifying AI Inference Deployments for Trillion Parameter Large Language Models — NVIDIA Technical Blog — developer.nvidia.com,” https://developer.nvidia.com/blog/demystifying-ai-inference-deployments-for-trillion-parameter-large-language-models/ , 2024, [Accessed 22-11-2024]
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Groq, “Groqracktm compute cluster - product brief v1.7.” [Online]. Available: https://groq.com/wp-content/uploads/2024/08/GroqRack%E2%84%A2-Compute-Cluster-Product-Brief-v1.7.pdf
2024
Closest in time.
“NVIDIA Blackwell Platform Arrives to Power a New Era of Computing — nvidianews.nvidia.com,” https://nvidianews.nvidia.com/news/nvidia-blackwell-platform-arrives-to-power-a-new-era-of-computing , 2024, [Accessed 20-11-2024]
2024
Closest in time.
“Nvidia h100 tensor core gpu datasheet,” https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet , 2023, (Accessed on 11/22/2024)
2024
Closest in time.
“hpc-datasheet-sc23-h200-datasheet-3002446.pdf,” https://nvdam.widen.net/s/nb5zzzsjdf/hpc-datasheet-sc23-h200-datasheet-3002446 , 2024, (Accessed on 11/22/2024)
2024
Closest in time.
“Aws graviton4 - annapurna labs (amazon) - wikichip,” https://en.wikichip.org/wiki/annapurna_labs/graviton/graviton4 , 2024, (Accessed on 11/22/2024)
2024
Closest in time.
“Prod_brief_qcom_cloud_ai_100_ultra.pdf,” https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Prod_Brief_QCOM_Cloud_AI_100_Ultra.pdf , 2021, (Accessed on 11/22/2024)
2024
Closest in time.
“Hotchips2020_ml_inference_baidu_kunlun_v5.pdf,” https://www.hc32.hotchips.org/assets/program/conference/day2/HotChips2020_ML_Inference_Baidu_Kunlun_v5.pdf , 2020, (Accessed on 11/22/2024)
2024
Closest in time.
“Optical interconnects set to achieve bandwidth benchmark — features — jul 2018 — photonics spectra,” https://www.photonics.com/Articles/Optical_Interconnects_Set_to_Achieve_Bandwidth/a63495 , 2018, (Accessed on 11/22/2024)
2024
Closest in time.
“solution-overview-gtcspring24-quantum-x800-3175164.pdf,” https://nvdam.widen.net/s/hbp8zz7fvt/solution-overview-gtcspring24-quantum-x800-3175164 , 2024, (Accessed on 11/23/2024)
2024
Closest in time.
“Ib_intro_wp_190.pdf,” https://network.nvidia.com/pdf/whitepapers/IB_Intro_WP_190.pdf , 2023, (Accessed on 11/23/2024)
2024
Closest in time.
“Hyperion-research-special-analysis-ualink-group-created-for-accelerator-interconnect-july-2024.pdf,” chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://hyperionresearch.com/wp-content/uploads/2024/07/Hyperion-Research-Special-Analysis-UALink-Group-Created-for-Accelerator-Interconnect-July-2024.pdf , 2024, (Accessed on 11/23/2024)
2024
Closest in time.
F. team, “Accelerating self-attentions for llm serving with flashinfer.” [Online]. Available: https://flashinfer.ai/2024/02/02/introduce-flashinfer.html
2024
Closest in time.
2024
Closest in time.
S. Hsia, A. Golden, B. Acun, N. Ardalani, Z. DeVito, G.-Y. Wei, D. Brooks, and C.-J. Wu, “Mad-max beyond single-node: Enabling large machine learning model acceleration on distributed systems,” in 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2024, pp. 818–833
2024
Closest in time.
“Claude 3.7 sonnet and claude code \ anthropic,” [Online; accessed 2025-04-10]. [Online]. Available: https://www.anthropic.com/news/claude-3-7-sonnet
2025
Closest in time.
“The llama 4 herd: The beginning of a new era of natively multimodal ai innovation,” [Online; accessed 2025-04-10]. [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/
2025
Closest in time.
“Etched is making the biggest bet in ai,” 2024, (Accessed on 2/21/2025). [Online]. Available: https://www.etched.com/announcing-etched
2025
Closest in time.
C. Trueman, “Photonic computing company lightmatter valued at $4.4bn in $400m funding round - dcd,” 10 2024, [Online; accessed 2025-04-09]. [Online]. Available: https://www.datacenterdynamics.com/en/news/photonic-computing-company-lightmatter-achieves-44bn-valuation-from-400m-series-d-funding-round/
2025
Closest in time.
“Sambaflow developer documentation :: Sambanova documentation,” [Online; accessed 2025-02-21]. [Online]. Available: https://docs.sambanova.ai/developer/latest/index.html
2025
Closest in time.
AMAX, “Comparing nvidia blackwell configurations,” 3 2024, [Online; accessed 2025-02-18]. [Online]. Available: https://www.amax.com/comparing-nvidia-blackwell-configurations/
2025
Closest in time.