Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Cited alongside, same era.
Learning to compose domain-specific transformations for data augmentation
Ratner, A., Ehrenberg, H., Hussain, Z., Dunnmon, J., and Re, C · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Hongler, C., and Gabriel, F · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2018
Cited alongside, same era.
Kernel machines that adapt to gpus for effective large batch training, 2018
Ma, S. and Belkin, M · 2018
Cited alongside, same era.
myrtle.ai, 2018
Page, D · 2018
Cited alongside, same era.
numpywren: Serverless linear algebra
Original
Shankar, V., Krauth, K., Pu, Q., Jonas, E., Venkataraman, S., Stoica, I., Recht, B., and Ragan-Kelley, J · 2018
Cited alongside, same era.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Original
Vasilache, N., Zinenko, O., Theodoridis, T., Goyal, P., DeVito, Z., Moses, W. S., Verdoolaege, S., Adams, A., and Cohen, A · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A
Cited in the paper.