Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Later among the works it cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Later among the works it cites.
Configurable clouds
Adrian M Caulfield, Eric S Chung, Andrew Putnam, Hari Angepat, Daniel Firestone, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, et al · 2017
Later among the works it cites.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
Christopher De Sa, Matthew Feldman, Christopher Ré, and Kunle Olukotun · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Later among the works it cites.
TensorQuant: A simulation toolbox for deep neural network quantization
Dominik Marek Loroch, Franz-Josef Pfreundt, Norbert Wehn, and Janis Keuper · 2017
Later among the works it cites.
Mixed precision training
Original
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2017
Later among the works it cites.
The ZipML framework for training models with end-to-end low precision: The cans, the cannots, and a little bit of deep learning
H Zhang, J Li, K Kara, D Alistarh, J Liu, and C Zhang · 2017
Later among the works it cites.
Microsoft unveils Project Brainwave for real-time ai
Doug Burger · 2018
Closest in time.
Transparent model distillation
Original
Sarah Tan, Rich Caruana, Giles Hooker, and Albert Gordo · 2018
Closest in time.