Fetching the paper…
Reading the bibliography…
Checkpoints play an important role in training long running machine learning (ML) models.
Distributed snapshots: Determining global states of distributed systems
K Mani Chandy and Leslie Lamport · 1985
Earlier work this paper cites.
Checkpointing and recovery rollback for distributed systems
R Koo and S Toueg · 1987
Earlier work this paper cites.
Job scheduling under the portable batch system
Robert L Henderson · 1995
Earlier work this paper cites.
On checkpoint latency
Nitin H Vaidya · 1995
Earlier work this paper cites.
An overview of checkpointing in uniprocessor and distributed systems, focusing on implementation and performance
James S Plank · 1997
Earlier work this paper cites.
System-level fault-tolerance in large-scale parallel machines with buffered coscheduling
Fabrizio Petrini, Kei Davis, and José Carlos Sancho · 2004
Earlier work this paper cites.
Modeling coordinated checkpointing for large-scale supercomputers
Long Wang, Karthik Pattabiraman, Zbigniew Kalbarczyk, Ravishankar K Iyer, Lawrence Votta, Christopher Vick, and Alan Wood · 2005
Earlier work this paper cites.
Smaller and faster data compression with zstandard
Yann Collet and Chip Turner · 2006
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Design, modeling, and evaluation of a scalable multi-level checkpointing system
Adam Moody, Greg Bronevetsky, Kathryn Mohror, and Bronis R De Supinski · 2010
Earlier work this paper cites.
The k-means clustering technique: General considerations and implementation in mathematica
Laurence Morissette and Sylvain Chartier · 2013
Earlier work this paper cites.
Optimization of multi-level checkpoint model for large scale hpc applications
Sheng Di, Mohamed Slim Bouguerra, Leonardo Bautista-Gomez, and Franck Cappello · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Bistro: Scheduling data-parallel jobs against live production systems
Andrey Goder, Alexey Spiridonov, and Yin Wang · 2015
Cited alongside, same era.
Revisiting distributed synchronous sgd
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Deep neural networks for youtube recommendations
Paul Covington, Jay Adams, and Emre Sargin · 2016
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Cited alongside, same era.
The netflix recommender system: Algorithms, business value, and innovation
Carlos A. Gomez-Uribe and Neil Hunt · 2016
Cited alongside, same era.
Fixed point quantization of deep convolutional networks
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016
Cited alongside, same era.
Lq-nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua · 2018
Later among the works it cites.
Bandana: Using non-volatile memory for storing deep learning models
Assaf Eisenman, Maxim Naumov, Darryl Gardner, Misha Smelyanskiy, Sergey Pupyrev, Kim M. Hazelwood, Asaf Cidon, and Sachin Katti · 2019
Later among the works it cites.
Post-training 4-bit quantization on embedding tables
Hui Guan, Andrey Malevich, Jiyan Yang, Jongsoo Park, and Hector Yuen · 2019
Later among the works it cites.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G Azzolini, et al · 2019
Later among the works it cites.
Haq: Hardware-aware automated quantization with mixed precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Two decades of recommender systems at amazon. com
Brent Smith and Greg Linden · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Adaptive quantization for deep neural network
Yiren Zhou, Seyed-Mohsen Moosavi-Dezfooli, Ngai-Man Cheung, and Pascal Frossard · 2017
Cited alongside, same era.
Billion-scale commodity embedding for e-commerce recommendation in alibaba
Jizhe Wang, Pipei Huang, Huan Zhao, Zhibo Zhang, Binqiang Zhao, and Dik Lun Lee · 2018
Cited alongside, same era.
Aibox: Ctr prediction model training on a single node
Weijie Zhao, Jingyuan Zhang, Deping Xie, Yulei Qian, Ronglai Jia, and Ping Li · 2019
Later among the works it cites.
Improving recommendation quality in google drive
Suming J Chen, Zhen Qin, Zac Wilson, Brian Calaci, Michael Rose, Ryan Evans, Sean Abraham, Donald Metzler, Sandeep Tata, and Michael Colagrosso · 2020
Closest in time.
The architectural implications of facebook’s dnn-based personalized recommendation
Udit Gupta, Carole-Jean Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim Hazelwood, Mark Hempstead, Bill Jia, et al · 2020
Closest in time.
Cross-stack workload characterization of deep recommendation systems
Samuel Hsia, Udit Gupta, Mark Wilkening, Carole-Jean Wu, Gu-Yeon Wei, and David Brooks · 2020
Closest in time.
Deep learning training in facebook data centers: Design of scale-up and scale-out systems, 2020
Maxim Naumov, John Kim, Dheevatsa Mudigere, Srinivas Sridharan, Xiaodong Wang, Whitney Zhao, Serhat Yilmaz, Changkyu Kim, Hector Yuen, Mustafa Ozdal, Krishnakumar Nair, Isabel Gao, Bor-Yiing Su, Jiyan Yang, and Mikhail Smelyanskiy · 2020
Closest in time.
Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models
Bogdan Nicolae, Jiali Li, Justin Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello · 2020
Closest in time.
Nvidia hgx2 datasheet
Nvidia · 2021
Closest in time.