Parameter-efficient transfer learning for nlp, 2019
Original
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 1902
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension, 2019
Original
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1904
Earlier work this paper cites.
Neural tangent kernels, transportation mappings, and universal approximation, 2019
Original
Ziwei Ji, Matus Telgarsky, and Ruicheng Xian · 1910
Earlier work this paper cites.
Topics in Matrix Analysis
Roger A. Horn and Charles R. Johnson · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A.R. Barron · 1993
Earlier work this paper cites.
Almost-sure identifiability of multidimensional harmonic retrieval
Tao Jiang, N.D. Sidiropoulos, and J.M.F. ten Berge · 2001
Earlier work this paper cites.
The zero set of a polynomial
Richard Caron and Tim Traynor · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Hadamard, khatri-rao, kronecker and other matrix products
Shuangzhe Liu and OTZ Trenkler · 2008
Earlier work this paper cites.
Uniform approximation of functions with random bases
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
A convergence analysis of log-linear training
Simon Wiesler and Hermann Ney · 2011
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.