2018

Hessian-based Analysis of Large Batch Training and Robustness to Adversaries

Yao, Zhewei, Gholami, Amir, Lei, Qi et al.

Understand

Large batch size training of Neural Networks has been shown to incur accuracy loss when trained with the current methods.

  • The exact underlying reasons for this are still not completely understood.
  • Here, we study large batch size training through the lens of the Hessian operator and robust optimization.
  • In particular, we perform a Hessian based study to analyze exactly how the landscape of the loss function changes when training with large batch size.

Reading the bibliography…