Fetching the paper…

Deep linear networks for regression are implicitly regularized towards flat minima · Around