“Wide neural networks of any depth evolve as linear models under gradient descent”, 2019
Original
Jaehoon Lee et al · 1902
Earlier work this paper cites.
“Selection via proxy: Efficient data selection for deep learning”, 2019
Original
Cody Coleman et al · 1906
Earlier work this paper cites.
“Coresets for data-efficient training of machine learning models”
Original
Baharan Mirzasoleiman, Jeff Bilmes and Jure Leskovec · 1906
Earlier work this paper cites.
“Relatif: Identifying explanatory training samples via relative influence”
Elnaz Barshan, Marc-Etienne Brunet and Gintare Dziugaite · 1909
Earlier work this paper cites.
“Harnessing the power of infinitely wide deep nets on small-data tasks”, 2019
Original
Sanjeev Arora et al · 1910
Earlier work this paper cites.
“Linear mode connectivity and the lottery ticket hypothesis”
Original
Jonathan Frankle, Gintare Dziugaite, Daniel Roy and Michael Carbin · 1912
Earlier work this paper cites.
“Submodularity in data subset selection and active learning”
Kai Wei, Rishabh Iyer and Jeff Bilmes · 1963
Earlier work this paper cites.
“The large learning rate phase of deep learning: the catapult mechanism”, 2020
Original
Aitor Lewkowycz et al · 2003
Earlier work this paper cites.
“Core vector machines: Fast SVM training on very large data sets.”
Ivor Tsang, James Kwok, Pak-Ming Cheung and Nello Cristianini · 2005
Earlier work this paper cites.
“Geometric approximation via coresets”
Pankaj Agarwal, Sariel Har-Peled and Kasturi Varadarajan · 2005
Earlier work this paper cites.
“Smaller coresets for k-median and k-means clustering”
Sariel Har-Peled and Akash Kushal · 2007
Earlier work this paper cites.
“What neural networks memorize and why: Discovering the long tail via influence estimation”, 2020
Original
Vitaly Feldman and Chiyuan Zhang · 2008
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”
Alex Krizhevsky, Vinod Nair and Geoffrey Hinton · 2009
Earlier work this paper cites.