Fetching the paper…
Reading the bibliography…
Emerging AI accelerators increasingly adopt wafer-scale manufacturing technologies, integrating hundreds of thousands of AI cores in a mesh architecture with large distributed on-chip memory (tens of GB in total) and ultra-high on-chip memory bandwidth (tens of PB/s).
A cellular computer to implement the kalman filter algorithm
Lynn Elliot Cannon · 1969
Earlier work this paper cites.
Parallel matrix transpose algorithms on distributed memory concurrent computers
Jaeyoung Choi, Jack J. Dongarra, and David W. Walker · 1995
Earlier work this paper cites.
SUMMA: scalable universal matrix multiplication algorithm
R. A. Van De Geijn and J. Watts · 1997
Earlier work this paper cites.
Multiple Si layer ICs: motivation, performance analysis, and design implications
Shukri J. Souri, Kaustav Banerjee, Amit Mehrotra, and Krishna C. Saraswat · 2000
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, et al · 2016
Earlier work this paper cites.
Tensor processing units for machine learning: An introduction
Norman Jouppi, Cliff Young, et al · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, et al · 2017
Earlier work this paper cites.
XLA: Optimizing TensorFlow for high performance
J. Rock et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
TVM: An automated end-to-end optimization stack for deep learning
T. Chen et al · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Rammer: Enabling holistic deep learning compiler optimizations with rTasks
Lingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue, Youshan Miao, Wei Cui, Wenxiang Hu, Fan Yang, Lintao Zhang, and Lidong Zhou · 2020
Earlier work this paper cites.
Ansor: A compiler stack for auto-tuning tensor programs
Y. Zhao et al · 2020
Earlier work this paper cites.
FlexTensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng · 2020
Earlier work this paper cites.
Chiplet/interposer co-design for power delivery network optimization in heterogeneous 2.5-d ICs
Jinwoo Kim, Venkata Chaitanya Krishna Chekuri, Nael Mizanur Rahman, Majid Ahadi Dolatsara, Hakki Mert Torun, Madhavan Swaminathan, Saibal Mukhopadhyay, and Sung Kyu Lim · 2021
Earlier work this paper cites.
TENET: A framework for modeling tensor dataflow based on relation-centric notation
Liqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia, Zizhang Luo, Jieming Yin, Jason Cong, and Yun Liang · 2021
Earlier work this paper cites.
Efficient large-scale language model training on GPU clusters using Megatron-LM
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
Earlier work this paper cites.
DOJO: The microarchitecture of Tesla’s exa-scale computer
Emil Talpes, Douglas Williams, and Debjit Das Sarma · 2022
Earlier work this paper cites.
Application defined on-chip networks for heterogeneous chiplets: An implementation perspective
Tianqi Wang, Fan Feng, Shaolin Xiang, Qi Li, and Jing Xia · 2022
Earlier work this paper cites.
DISTAL: the distributed tensor algebra compiler
Rohan Yadav, Alex Aiken, and Fredrik Kjolstad · 2022
Cited alongside, same era.
Alpa: Automating inter-and intra-operator parallelism for distributed deep learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al · 2022
Cited alongside, same era.
Exploring TensorRT to improve real-time inference for deep learning
Yuxiao Zhou and Kecheng Yang · 2022
Cited alongside, same era.
ROLLER: Fast and efficient tensor compilation for deep learning
Hongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke, Haoyu Li, Chen Zhang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Wei Cui, Fan Yang, Mao Yang, Lidong Zhou, Asaf Cidon, and Gennady Pekhimenko · 2022
Cited alongside, same era.
AMD optimizes EPYC memory with NUMA
Advanced Micro Devices · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
TSMC bets big on advanced packaging, 2023
Mark LaPedus · 2024
Later among the works it cites.
ReSA: Reconfigurable systolic array for multiple tiny DNN tensors
Ching-Jui Lee and Tsung Tai Yeh · 2024
Later among the works it cites.
Scaling deep learning computation over the inter-core connected intelligence processor with T10
Yiqi Liu, Yuqi Xue, Yu Cheng, Lingxiao Ma, Ziming Miao, Jilong Xue, and Jian Huang · 2024
Later among the works it cites.
Near-optimal wafer-scale reduce
Piotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson, Daniele De Sensi, and Torsten Hoefler · 2024
Later among the works it cites.
An electrical-thermal co-simulation model of chiplet heterogeneous integration systems
Xiaoning Ma, Qinzhi Xu, Chenghan Wang, He Cao, Jianyun Liu, Daoqing Zhang, and Zhiqiang Li · 2024
Later among the works it cites.
Introducing MTIA: Meta’s next-generation training and inference accelerator for AI, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Flashdecoding++: Faster large language model inference on GPUs
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, Kangdi Chen, Hanyu Dong, and Yu Wang · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning
Sean Lie · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean · 2023
Cited alongside, same era.
Welder: Scheduling deep learning memory access via tile-graph
Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou · 2023
Cited alongside, same era.
Meta AI · 2024
Later among the works it cites.
Azure Maia: For the era of AI from silicon to software to systems, 2023
Microsoft Azure · 2024
Later among the works it cites.
Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation
Lei Wang, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi, Ningxin Zheng, Ziming Miao, Fan Yang, Ting Cao, et al · 2024
Later among the works it cites.
Static random-access memory, 2024
Wikipedia contributors · 2024
Later among the works it cites.
Wafer-scale integration, 2024
Wikipedia contributors · 2024
Later among the works it cites.
LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2024
Later among the works it cites.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Later among the works it cites.
100× defect tolerance: How cerebras solved the yield problem, 2022
Cerebras Systems · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Openai o3 and o4-mini system card
OpenAI · 2025
Closest in time.
Cerebras and g42 break ground on condor galaxy 3, an 8 exaflops ai supercomputer
Cerebras Systems · 2025
Closest in time.
Cerebras powers perplexity sonar with industry’s fastest ai inference
Cerebras Systems · 2025
Closest in time.
Cerebras brings instant inference to mistral le chat
James Wang · 2025
Closest in time.