Fetching the paper…
Reading the bibliography…
The rapid progress in Deep Learning (DL) and Large Language Models (LLMs) has exponentially increased demands of computational power and bandwidth.
1909
Earlier work this paper cites.
C. E. Leiserson, “Fat-trees: Universal networks for hardware-efficient supercomputing,” IEEE Transactions on Computers , vol. C-34, no. 10, pp. 892–901, Oct 1985
1985
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation , vol. 3, no. 1, pp. 79–87, 1991
1991
Earlier work this paper cites.
M. I. Jordan and R. A. Jacobs, “Hierarchical mixtures of experts and the em algorithm,” Neural computation , vol. 6, no. 2, pp. 181–214, 1994
1994
Earlier work this paper cites.
R. R. Schaller, “Moore’s law: past, present and future,” IEEE spectrum , vol. 34, no. 6, pp. 52–59, 1997
1997
Earlier work this paper cites.
R. Budruk, D. Anderson, and E. Solari, PCI Express System Architecture . Pearson Education, 2003
2003
Earlier work this paper cites.
S.-A. Reinemo, T. Skeie, T. Sodring, O. Lysne, and O. Trudbakken, “An overview of qos capabilities in infiniband, advanced switching interconnect, and ethernet,” IEEE Communications Magazine , vol. 44, no. 7, pp. 32–38, 2006
2006
Earlier work this paper cites.
M. Scharf and S. Kiesel, “Nxg03-5: Head-of-line blocking in tcp and sctp: Analysis and measurements,” in IEEE Globecom 2006 , 2006, pp. 1–5
2006
Earlier work this paper cites.
S. Yan, G. Min, and I. Awan, “Performance analysis of credit-based flow control in infiniband interconnection networks,” Journal of Interconnection Networks , vol. 07, no. 04, pp. 535–548, 2006. [Online]. Available: https://doi.org/10.1142/S0219265906001843
2006
Earlier work this paper cites.
J. Kim, W. J. Dally, S. Scott, and D. Abts, “Technology-driven, highly-scalable dragonfly topology,” in 2008 International Symposium on Computer Architecture , 2008, pp. 77–88
2008
Earlier work this paper cites.
H. Subramoni, P. Lai, M. Luo, and D. K. Panda, “Rdma over ethernet — a preliminary study,” in 2009 IEEE International Conference on Cluster Computing and Workshops , 2009, pp. 1–9
2009
Earlier work this paper cites.
P. Sanders, J. Speck, and J. Träff, “Two-tree algorithms for full bandwidth broadcast, reduction and scan,” Parallel Computing , vol. 35, pp. 581–594, 12 2009
2009
Earlier work this paper cites.
J. Terrace and M. J. Freedman, “Object storage on CRAQ: High-Throughput chain replication for Read-Mostly workloads,” in 2009 USENIX Annual Technical Conference (USENIX ATC 09) . San Diego, CA: USENIX Association, Jun. 2009. [Online]. Available: https://www.usenix.org/conference/usenix-09/object-storage-craq-high-throughput-chain-replication-read-mostly-workloads
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
E. B. Nightingale, J. Elson, J. Fan, O. Hofmann, J. Howell, and Y. Suzue, “Flat datacenter storage,” in Proceedings of the 10th USENIX Conference on Operating Systems Design and Implementation , ser. OSDI’12. USA: USENIX Association, 2012, p. 1–15
2012
Earlier work this paper cites.
R. Shi, S. Potluri, K. Hamidouche, J. Perkins, M. Li, D. Rossetti, and D. K. D. K. Panda, “Designing efficient small message transfer mechanism for inter-node mpi communication on infiniband gpu clusters,” in 2014 21st International Conference on High Performance Computing (HiPC) , 2014, pp. 1–10
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
S. Liu and W. Deng, “Very deep convolutional neural network based image classification using small training sample size,” in 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR) , 2015, pp. 730–734
2015
Earlier work this paper cites.
“Priority flow control : Build reliable layer 2 infrastructure,” 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:42645413
2015
Earlier work this paper cites.
Y. Zhu, H. Eran, D. Firestone, C. Guo, M. Lipshteyn, Y. Liron, J. Padhye, S. Raindel, M. H. Yahia, and M. Zhang, “Congestion control for large-scale rdma deployments,” SIGCOMM Comput. Commun. Rev. , vol. 45, no. 4, p. 523–536, aug 2015. [Online]. Available: https://doi.org/10.1145/2829988.2787484
2015
Earlier work this paper cites.
R. Mittal, V. T. Lam, N. Dukkipati, E. R. Blem, H. M. G. Wassel, M. Ghobadi, A. Vahdat, Y. Wang, D. Wetherall, and D. Zats, “Timely: Rtt-based congestion control for the datacenter,” Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication , 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:9676894
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
H. Fu, J. Liao, J. Yang, L. Wang, Z. Song, X. Huang, C. Yang, W. Xue, F. Liu, F. Qiao, W. Zhao, X. Yin, C. Hou, C. Zhang, W. Ge, J. Zhang, Y. Wang, C. Zhou, and G. Yang, “The Sunway TaihuLight supercomputer: system and applications,” Science China Information Sciences , vol. 59, no. 7, p. 072001, Jul. 2016. [Online]. Available: http://link.springer.com/10.1007/s11432-016-5588-7
2016
Earlier work this paper cites.
P. Foundation, “Tensors and dynamic neural networks in python with strong gpu acceleration,” 2016. [Online]. Available: https://github.com/pytorch/pytorch
2016
Earlier work this paper cites.
NVIDIA, “Nvidia collective communications library (nccl): Optimized primitives for collective multi-gpu communication,” 2017. [Online]. Available: https://github.com/NVIDIA/nccl
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Dongarra, “REPORT ON THE TIANHE-2A SYSTEM,” 2017
2017
Earlier work this paper cites.
D. Stanzione, B. Barth, N. Gaffney, K. Gaither, C. Hempel, T. Minyard, S. Mehringer, E. Wernert, H. Tufo, D. Panda, and P. Teller, “Stampede 2: The evolution of an xsede supercomputer,” in Proceedings of the Practice and Experience in Advanced Research Computing 2017 on Sustainability, Success and Impact , ser. PEARC ’17. New York, NY, USA: Association for Computing Machinery, 2017. [Online]. Available: https://doi.org/10.1145/3093338.3093385
2017
Earlier work this paper cites.
Z. Liran, H. David, and M. Barbara, “Wekafs architecture white paper,” 2021. [Online]. Available: https://www.weka.io/wp-content/uploads/files/2017/12/Architectural_WhitePaper-W02R6WP201812-1.pdf
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel et al. , “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science , vol. 362, no. 6419, pp. 1140–1144, 2018
2018
Earlier work this paper cites.
H. Frank and B. Sven, “Wekafs architecture white paper,” 2018. [Online]. Available: https://www.beegfs.io/docs/whitepapers/Introduction_to_BeeGFS_by_ThinkParQ.pdf
2018
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , ser. SOSP ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 1–15. [Online]. Available: https://doi.org/10.1145/3341301.3359646
2019
Cited alongside, same era.
H. Liao, J. Tu, J. Xia, and X. Zhou, “Davinci: A scalable architecture for neural network computing,” in 2019 IEEE Hot Chips 31 Symposium (HCS) , 2019, pp. 1–44
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics , 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:52967399
2019
Cited alongside, same era.
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y. He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,” in Proceedings of the 37th ACM International Conference on Supercomputing , ser. ICS ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 203–214. [Online]. Available: https://doi.org/10.1145/3577193.3593704
2023
Later among the works it cites.
C. Hwang, W. Cui, Y. Xiong, Z. Yang, Z. Liu, H. Hu, Z. Wang, R. Salas, J. Jose, P. Ram et al. , “Tutel: Adaptive mixture-of-experts at scale,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
HFAiLab, “Hai platform: A high-performance deep learning training platform with task-level gpu compute time-sharing scheduling,” 2023. [Online]. Available: https://github.com/HFAiLab/hai-platform
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Cited alongside, same era.
Y. Li, R. Miao, H. H. Liu, Y. Zhuang, F. Feng, L. Tang, Z. Cao, M. Zhang, F. Kelly, M. Alizadeh, and M. Yu, “Hpcc: high precision congestion control,” in Proceedings of the ACM Special Interest Group on Data Communication , ser. SIGCOMM ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 44–58. [Online]. Available: https://doi.org/10.1145/3341302.3342085
2019
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Cited alongside, same era.
A. W. Senior, R. Evans, J. Jumper, J. Kirkpatrick, L. Sifre, T. Green, C. Qin, A. Žídek, A. W. Nelson, A. Bridgland et al. , “Improved protein structure prediction using potentials from deep learning,” Nature , vol. 577, no. 7792, pp. 706–710, 2020
2020
Cited alongside, same era.
L. Floridi and M. Chiriatti, “Gpt-3: Its nature, scope, limits, and consequences,” Minds and Machines , vol. 30, pp. 681–694, 2020
2020
Cited alongside, same era.
T. Shimizu, “Supercomputer Fugaku: Co-designed with application developers/researchers,” in 2020 IEEE Asian Solid-State Circuits Conference (A-SSCC) . Hiroshima, Japan: IEEE, Nov. 2020, pp. 1–4. [Online]. Available: https://ieeexplore.ieee.org/document/9336127/
2020
Cited alongside, same era.
C. B. Stunkel, R. L. Graham, G. Shainer, M. Kagan, S. S. Sharkawi, B. Rosenburg, and G. A. Chochia, “The high-speed networks of the Summit and Sierra supercomputers,” IBM Journal of Research and Development , vol. 64, no. 3/4, pp. 3:1–3:10, May 2020. [Online]. Available: https://ieeexplore.ieee.org/document/8961159/
2020
Cited alongside, same era.
2023
Later among the works it cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
V. A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” in Proceedings of the Sixth Conference on Machine Learning and Systems, MLSys 2023, Miami, FL, USA, June 4-8, 2023 , D. Song, M. Carbin, and T. C. 0001, Eds. mlsys.org, 2023. [Online]. Available: https://proceedings.mlsys.org/paper_files/paper/2023/hash/80083951326cf5b35e5100260d64ed81-Abstract-mlsys2023.html
2023
Later among the works it cites.
A. N. Laboratory, “Argonne’s aurora supercomputer,” 2023. [Online]. Available: https://www.alcf.anl.gov/aurora
2023
Later among the works it cites.
2023
Later among the works it cites.
NVIDIA, “Nvidia announces dgx h100 systems – world’s most advanced enterprise ai infrastructure,” 2023. [Online]. Available: https://nvidianews.nvidia.com/news/nvidia-announces-dgx-h100-systems-worlds-most-advanced-enterprise-ai-infrastructure
2023
Later among the works it cites.
N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. A. Patterson, “TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture . Orlando FL USA: ACM, Jun. 2023, pp. 1–14. [Online]. Available: https://dl.acm.org/doi/10.1145/3579371.3589350
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Talpes, D. D. Sarma, D. Williams, S. Arora, T. Kunjan, B. Floering, A. Jalote, C. Hsiong, C. Poorna, V. Samant, J. Sicilia, A. K. Nivarti, R. Ramachandran, T. Fischer, B. Herzberg, B. McGee, G. Venkataramanan, and P. Banon, “The microarchitecture of dojo, tesla’s exa-scale computer,” IEEE Micro , vol. 43, no. 3, pp. 31–39, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Hennecke, “Understanding daos storage performance scalability,” in Proceedings of the HPC Asia 2023 Workshops , ser. HPCAsia ’23 Workshops. New York, NY, USA: Association for Computing Machinery, 2023, p. 1–14. [Online]. Available: https://doi.org/10.1145/3581576.3581577
2023
Later among the works it cites.
K. Liu, Z. Jiang, J. Zhang, H. Wei, X. Zhong, L. Tan, T. Pan, and T. Huang, “Hostping: Diagnosing intra-host network bottlenecks in RDMA servers,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . Boston, MA: USENIX Association, Apr. 2023, pp. 15–29. [Online]. Available: https://www.usenix.org/conference/nsdi23/presentation/liu-kefei
2023
Later among the works it cites.
Y. He, M. Hutton, S. Chan, R. De Gruijl, R. Govindaraju, N. Patil, and Y. Li, “Understanding and Mitigating Hardware Failures in Deep Learning Training Systems,” in Proceedings of the 50th Annual International Symposium on Computer Architecture . Orlando FL USA: ACM, Jun. 2023, pp. 1–16. [Online]. Available: https://dl.acm.org/doi/10.1145/3579371.3589105
2023
Later among the works it cites.
S. Wang, G. Zhang, J. Wei, Y. Wang, J. Wu, and Q. Luo, “Understanding Silent Data Corruptions in a Large Production CPU Population,” in Proceedings of the 29th Symposium on Operating Systems Principles . Koblenz Germany: ACM, Oct. 2023, pp. 216–230. [Online]. Available: https://dl.acm.org/doi/10.1145/3600006.3613149
2023
Later among the works it cites.
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh, “Video generation models as world simulators,” 2024. [Online]. Available: https://openai.com/research/video-generation-models-as-world-simulators
2024
Closest in time.
2024
Closest in time.
Z. Sun, H. Cao, Y. Wang, G. Feng, S. Chen, H. Wang, and W. Chen, “Adapipe: Optimizing pipeline parallelism with adaptive recomputation and partitioning,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , ser. ASPLOS ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 86–100. [Online]. Available: https://doi.org/10.1145/3620666.3651359
2024
Closest in time.
X. Hou, Y. Yuan, S. Ma, R. Xu, B. Wang, T. Li, W. Jiang, L. Wu, and J. Zhang, “Optimizing the parallelism of communication and computation in distributed training platform,” in Algorithms and Architectures for Parallel Processing: 23rd International Conference, ICA3PP 2023, Tianjin, China, October 20–22, 2023, Proceedings, Part I . Berlin, Heidelberg: Springer-Verlag, 2024, p. 340–359. [Online]. Available: https://doi.org/10.1007/978-981-97-0834-5_20
2024
Closest in time.
A. Gangidi, R. Miao, S. Zheng, S. J. Bondu, G. Goes, H. Morsy, R. Puri, M. Riftadi, A. J. Shetty, J. Yang, S. Zhang, M. J. Fernandez, S. Gandham, and H. Zeng, “Rdma over ethernet for distributed training at meta scale,” in Proceedings of the ACM SIGCOMM 2024 Conference , ser. ACM SIGCOMM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 57–70. [Online]. Available: https://doi.org/10.1145/3651890.3672233
2024
Closest in time.
2024
Closest in time.
K. Qian, Y. Xi, J. Cao, J. Gao, Y. Xu, Y. Guan, B. Fu, X. Shi, F. Zhu, R. Miao, C. Wang, P. Wang, P. Zhang, X. Zeng, E. Ruan, Z. Yao, E. Zhai, and D. Cai, “Alibaba hpn: A data center network for large language model training,” in Proceedings of the ACM SIGCOMM 2024 Conference , ser. ACM SIGCOMM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 691–706. [Online]. Available: https://doi.org/10.1145/3651890.3672265
2024
Closest in time.
DeepSeek-AI, “Deepseek api introduces context caching on disk, cutting prices by an order of magnitude,” 2024. [Online]. Available: https://platform.deepseek.com/api-docs/news/news0802
2024
Closest in time.
Q. Hu, Z. Ye, Z. Wang, G. Wang, M. Zhang, Q. Chen, P. Sun, D. Lin, X. Wang, Y. Luo et al. , “Characterization of large language model development in the datacenter,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 709–729
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.