Fetching the paper…

Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks · Around