Two models of double descent for weak features
Original
Belkin, M., Hsu, D., and Xu, J · 1903
Earlier work this paper cites.
A class of wasserstein metrics for probability distributions
Givens, C. R., Shortt, R. M., et al · 1984
Earlier work this paper cites.
Random matrix theory and wireless communications
Tulino, A. M., Verdú, S., et al · 2004
Earlier work this paper cites.
Expectation consistent approximate inference
Opper, M. and Winther, O · 2005
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Villani, C · 2008
Earlier work this paper cites.
Message-passing algorithms for compressed sensing
Donoho, D. L., Maleki, A., and Montanari, A · 2009
Earlier work this paper cites.
Message passing algorithms for compressed sensing
Donoho, D. L., Maleki, A., and Montanari, A · 2010
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Bayati, M. and Montanari, A · 2011
Earlier work this paper cites.
S-AMP: Approximate message passing for general matrix ensembles
Cakmak, B., Winther, O., and Fleury, B. H · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Earlier work this paper cites.
Expectation consistent approximate inference: Generalizations and convergence
Fletcher, A., Sahraee-Ardakan, M., Rangan, S., and Schniter, P · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Original
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.