2018

Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD

Dutta, Sanghamitra, Joshi, Gauri, Ghosh, Soumyadip et al.

Understand

Distributed Stochastic Gradient Descent (SGD) when run in a synchronous manner, suffers from delays in waiting for the slowest learners (stragglers).

  • Asynchronous methods can alleviate stragglers, but cause gradient staleness that can adversely affect convergence.
  • In this work we present a novel theoretical characterization of the speed-up offered by asynchronous methods by analyzing the trade-off between the error in the trained model and the actual training runtime (wallclock time).
  • The novelty in our work is that our runtime analysis considers random straggler delays, which helps us design and compare distributed SGD algorithms that strike a balance between stragglers and staleness.

Reading the bibliography…