Fantastic generalization measures and where to find them
Original
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Large scale learning of general visual representations for transfer
Original
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2019
Later among the works it cites.
DARTS: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Later among the works it cites.
Generalization bounds for deep convolutional neural networks
Original
Philip M Long and Hanie Sedghi · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning, 2019
Vaishnavh Nagarajan and J. Zico Kolter · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Later among the works it cites.
Improved sample complexities for deep neural networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Closest in time.
Small data, big decisions: Model selection in the small-data regime
Jörg Bornschein, Francesco Visin, and Simon Osindero · 2020
Closest in time.
Language models are few-shot learners
Original
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeff Wu, Heewoo Jun, Prafulla Dhariwal, David Luan, and Ilya Sutskever · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Original
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Nas-bench-201: Extending the scope of reproducible neural architecture search
Xuanyi Dong and Yi Yang · 2020
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Closest in time.
In search of robust measures of generalization
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar, Ethan Caballero, Linbo Wang, Ioannis Mitliagkas, and Daniel M Roy · 2020
Closest in time.
Threedworld: A platform for interactive multi-modal physical simulation
Original
Chuang Gan, Jeremy Schwartz, Seth Alter, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, et al · 2020
Closest in time.
Affinity and diversity: Quantifying mechanisms of data augmentation
Original
Raphael Gontijo-Lopes, Sylvia J Smullin, Ekin D Cubuk, and Ethan Dyer · 2020
Closest in time.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, St’efan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fern’andez del R’ıo, Mark Wiebe, Pearu Peterson, Pierre G’erard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Closest in time.
Denoising diffusion probabilistic models
Original
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Closest in time.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Original
Like Hui and Mikhail Belkin · 2020
Closest in time.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Distributional generalization: A new kind of generalization
Original
Preetum Nakkiran and Yamini Bansal · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Closest in time.
Towards learning convolutions from scratch
Original
Behnam Neyshabur · 2020
Closest in time.