Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Convex optimization
Stephen P Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
An improved estimate in the restricted isometry problem
Jean Bourgain · 2014
Earlier work this paper cites.
Sample-optimal Fourier sampling in any constant dimension
Piotr Indyk and Michael Kapralov · 2014
Earlier work this paper cites.
(nearly) sample-optimal sparse fourier transform
Piotr Indyk, Michael Kapralov, and Eric Price · 2014
Earlier work this paper cites.
Fourier-sparse interpolation without a frequency gap
Xue Chen, Daniel M Kane, Eric Price, and Zhao Song · 2016
Earlier work this paper cites.
Sparse Fourier transform in any constant dimension with nearly-optimal sample complexity in sublinear time
Michael Kapralov · 2016
Earlier work this paper cites.
The restricted isometry property of subsampled fourier matrices
Ishay Haviv and Oded Regev · 2017
Earlier work this paper cites.
Sample efficient estimation and recovery in sparse FFT via isolation on average
Michael Kapralov · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
(nearly) sample-optimal sparse fourier transform in any dimension; ripless and filterless
Vasileios Nakos, Zhao Song, and Zhengyu Wang · 2019
Earlier work this paper cites.
Matrix Theory: Optimization, Concentration and Algorithms
Zhao Song · 2019
Earlier work this paper cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Earlier work this paper cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Earlier work this paper cites.
Curve detectors
Nick Cammarata, Gabriel Goh, Shan Carter, Ludwig Schubert, Michael Petrov, and Chris Olah · 2020
Earlier work this paper cites.
Learning mixtures of linear regressions in subexponential time via fourier moments
Sitan Chen, Jerry Li, and Zhao Song · 2020
Earlier work this paper cites.
Learning parities with neural networks
Amit Daniely and Eran Malach · 2020
Earlier work this paper cites.
Multi-pass transformer for machine translation
Original
Peng Gao, Chiori Hori, Shijie Geng, Takaaki Hori, and Jonathan Le Roux · 2020
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Fully quantized transformer for machine translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh · 2020
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Original
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Earlier work this paper cites.
On the inductive bias of neural networks for learning read-once dnfs
Ido Bronstein, Alon Brutzkus, and Amir Globerson · 2022
Earlier work this paper cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Earlier work this paper cites.