Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) are the de facto standard for essential use cases, such as image classification, computer vision, and natural language processing.
Hadamard matrices and their applications
A Hedayat and Walter Dennis Wallis · 1978
Earlier work this paper cites.
Approximate nearest neighbors and the fast johnson-lindenstrauss transform
Nir Ailon and Bernard Chazelle · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Uncertainty Principles and Vector Quantization
Yurii Lyubarskii and Roman Vershynin · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
InfiniBand Trade Association. RoCE v2 Specification.
InfiniBand Trade Association · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2014
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Pingmesh: A large-scale system for data center network latency measurement and analysis
Chuanxiong Guo, Lihua Yuan, Dong Xiang, Yingnong Dang, Ray Huang, Dave Maltz, Zhaoyi Liu, Vin Wang, Bin Pang, Hua Chen, Zhi-Wei Lin, and Varugis Kurien · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Lossradar: Fast detection of lost packets in data center networks
Yuliang Li, Rui Miao, Changhoon Kim, and Minlan Yu · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Database meets deep learning: Challenges and opportunities
Wei Wang, Meihui Zhang, Gang Chen, H. V. Jagadish, Beng Chin Ooi, and Kian-Lee Tan · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Z. Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Distributed Mean Estimation With Limited Communication
Ananda Theertha Suresh, X Yu Felix, Sanjiv Kumar, and H Brendan McMahan · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
A survey on homomorphic encryption schemes: Theory and implementation
Abbas Acar, Hidayet Aksu, A. Selcuk Uluagac, and Mauro Conti · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A configurable cloud-scale dnn processor for real-time ai
Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Todd Massengill, Ming Liu, Daniel Lo, Shlomi Alkalay, Michael Haselman, Logan Adams, Mahdi Ghandi, Stephen Heil, Prerak Patel, Adam Sapek, Gabriel Weisz, Lisa Woods, Sitaram Lanka, Steven K. Reinhardt, Adrian M. Caulfield, Eric S. Chung, and Doug Burger · 2018
Earlier work this paper cites.
Integrated model, batch, and domain parallelism in training neural networks
Amir Gholami, Ariful Azad, Peter Jin, Kurt Keutzer, and Aydin Buluc · 2018
Earlier work this paper cites.
Randomized distributed mean estimation: Accuracy vs. communication
Jakub Konečnỳ and Peter Richtárik · 2018
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Sparsified sgd with memory
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Machine learning in compiler optimization
Zheng Wang and Michael O’Boyle · 2018
Cited alongside, same era.
Analysis of Large-Scale Multi-Tenant GPU clusters for DNN training workloads
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang · 2019
Cited alongside, same era.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Sip-ml: High-bandwidth optical network interconnects for machine learning training
Mehrdad Khani, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, and Eiman Ebrahimi · 2021
Later among the works it cites.
ATP: In-network aggregation for multi-tenant learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, and Michael Swift · 2021
Later among the works it cites.
On distributed adaptive optimization with gradient compression
Xiaoyun Li, Belhal Karimi, and Ping Li · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
Later among the works it cites.
Nuqsgd: Provably communication-efficient data-parallel sgd via nonuniform quantization
Ali Ramezani-Kebrya, Fartash Faghri, Ilya Markov, Vitalii Aksenov, Dan Alistarh, and Daniel M Roy · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Full-stack optimization for accelerating cnns using powers-of-two weights with fpga validation
Bradley McDanel, Sai Qian Zhang, H. T. Kung, and Xin Dong · 2019
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G Azzolini, et al · 2019
Cited alongside, same era.
A generic communication scheduler for distributed dnn training acceleration
Yanghua Peng, Yibo Zhu, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu, and Chuanxiong Guo · 2019
Cited alongside, same era.
When should the network be the computer?
Dan R. K. Ports and Jacob Nelson · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Scaling distributed machine learning with In-Network aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan Ports, and Peter Richtarik · 2021
Later among the works it cites.
Drive: One-bit distributed mean estimation
Shay Vargaftik, Ran Ben-Basat, Amit Portnoy, Gal Mendelson, Yaniv Ben-Itzhak, and Michael Mitzenmacher · 2021
Later among the works it cites.
On the utility of gradient compression in distributed training systems
Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman, and Dimitris Papailiopoulos · 2022
Later among the works it cites.
QUIC-FL: Quick Unbiased Compression for Federated Learning
Ran Ben Basat, Shay Vargaftik, Amit Portnoy, Gil Einziger, Yaniv Ben-Itzhak, and Michael Mitzenmacher · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Aaron Daniel Cohen, Adam Roberts, Alejandra Molina, Alena Butryna, Alicia Jin, Apoorv Kulshreshtha, Ben Hutchinson, Ben Zevenbergen, Blaise Hilary Aguera-Arcas, Chung ching Chang, Claire Cui, Cosmo Du, Daniel De Freitas Adiwardana, Dehao Chen, Dmitry (Dima) Lepikhin, Ed H. Chi, Erin Hoffman-John, Heng-Tze Cheng, Hongrae Lee, Igor Krivokon, James Qin, Jamie Hall, Joe Fenton, Johnny Soraker, Kathy Meier-Hellstern, Kristen Olson, Lora Mois Aroyo, Maarten Paul Bosma, Marc Joseph Pickett, Marcelo Amorim Menegali, Marian Croak, Mark Díaz, Matthew Lamm, Maxim Krikun, Meredith Ringel Morris, Noam Shazeer, Quoc V. Le, Rachel Bernstein, Ravi Rajakumar, Ray Kurzweil, Romal Thoppilan, Steven Zheng, Taylor Bos, Toju Duke, Tulsee Doshi, Vincent Y. Zhao, Vinodkumar Prabhakaran, Will Rusch, YaGuang Li, Yanping Huang, Yanqi Zhou, Yuanzhong Xu, and Zhifeng Chen · 2022
Later among the works it cites.
Kaja Gruntkowska, Alexander Tyurin, and Peter Richtárik · 2022
Later among the works it cites.
Preemptive switch memory usage to accelerate training jobs with shared in-network aggregation
Wang Hao, Qin Yuxuan, Lao ChonLam, Le Yanfang, Wu Wenfei, and Chen Kai · 2022
Later among the works it cites.
DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Later among the works it cites.
Compute trends across three eras of machine learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos · 2022
Later among the works it cites.
Eden: Communication-efficient and robust distributed mean estimation for federated learning
Shay Vargaftik, Ran Ben Basat, Amit Portnoy, Gal Mendelson, Yaniv Ben Itzhak, and Michael Mitzenmacher · 2022
Later among the works it cites.
Mlaas in the wild: Workload analysis and scheduling in large-scale heterogeneous gpu clusters
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding · 2022
Later among the works it cites.
Using trio: Juniper networks’ programmable chipset - for emerging in-network applications
Mingran Yang, Alex Baban, Valery Kugel, Jeff Libby, Scott Mackie, Swamy Sadashivaiah Renu Kananda, Chang-Hong Wu, and Manya Ghobadi · 2022
Later among the works it cites.
Unlocking the power of inline Floating-Point operations on programmable switches
Yifan Yuan, Omar Alama, Jiawei Fei, Jacob Nelson, Dan R. K. Ports, Amedeo Sapio, Marco Canini, and Nam Sung Kim · 2022
Later among the works it cites.
Docofl: Downlink compression for cross-device federated learning
Ron Dorfman, Shay Vargaftik, Yaniv Ben-Itzhak, and Kfir Yehuda Levy · 2023
Closest in time.
A generic service to provide in-network aggregation for key-value streams
Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu, and ChonLam Lao · 2023
Closest in time.
Janus: A unified distributed training framework for sparse mixture-of-experts models
Juncai Liu, Jessie Hui Wang, and Yimin Jiang · 2023
Closest in time.
Cassini: Network-aware job scheduling in machine learning clusters
Sudarsanan Rajasekaran, Manya Ghobadi, and Aditya Akella · 2023
Closest in time.
TopoOpt: Co-optimizing network topology and parallelization strategy for distributed training jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch · 2023
Closest in time.
Hi-speed dnn training with espresso: Unleashing the full potential of gradient compression with near-optimal usage strategies
Zhuang Wang, Haibin Lin, Yibo Zhu, and T. S. Eugene Ng · 2023
Closest in time.
Vitis AI
Xilinx · 2023
Closest in time.
Optimal and Near-Optimal Adaptive Vector Quantization
Ran Ben-Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, and Shay Vargaftik · 2024
Closest in time.