Fetching the paper…
Reading the bibliography…
This paper considers the Pointer Value Retrieval (PVR) benchmark introduced in [ZRKB21], where a 'reasoning' function acts on a string of digits to produce the label.
Weakly learning DNF and characterizing statistical query learning using Fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Dataset Shift in Machine Learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Vaughan · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning, 2014
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Analysis of Boolean Functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Implicit regularization in matrix factorization, 2017
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Earlier work this paper cites.
Exploring generalization in deep learning, 2017
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data, 2017
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Earlier work this paper cites.
Wasserstein distance guided representation learning for domain adaptation, 2017
Jian Shen, Yanru Qu, Weinan Zhang, and Yong Yu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach, 2018
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry, 2018
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks, 2018
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Training behavior of deep neural network in frequency domain, 2018
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization, 2019
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
On the learnability of deep random networks, 2019
Abhimanyu Das, Sreenivas Gollapudi, Ravi Kumar, and Rina Panigrahy · 2019
Cited alongside, same era.
Boolean functions with biased inputs: Approximation and noise sensitivity
Mohsen Heidari, S. Sandeep Pradhan, and Ramji Venkataramanan · 2019
Cited alongside, same era.
Explaining ai decisions using efficient methods for learning sparse boolean formulae
Susmit Jha, Tuhin Sahai, Vasumathi Raman, Alessandro Pinto, and Michael Francis · 2019
Implicit bias in deep linear classification: Initialization scale vs training accuracy, 2020
Edward Moroshko, Suriya Gunasekar, Blake Woodworth, Jason D. Lee, Nathan Srebro, and Daniel Soudry · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms, 2020
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
The staircase property: How hierarchical structure can guide deep learning, NeurIPS, 2021
Emmanuel Abbe, Enric Boix-Adsera, Matthew Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Later among the works it cites.
On the power of differentiable learning versus PAC and SQ learning
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon, and Nathan Srebro · 2021
Later among the works it cites.
Epistatic net allows the sparse spectral regularization of deep neural networks for inferring fitness functions
Amirali Aghazadeh, Hunter Nisonoff, Orhan Ocal, David H. Brookes, Yijie Huang, O. Ozan Koyluoglu, Jennifer Listgarten, and Kannan Ramchandran · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Characterizing the implicit bias via a primal-dual analysis, 2019
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks, 2019
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks, 2019
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2019
Cited alongside, same era.
Accuracy on the line: On the strong correlation between out-of-distribution and in-distribution generalization, 2021
John Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity, 2021
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Extending the wilds benchmark for unsupervised adaptation, 2021
Shiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao, Sang Michael Xie, Kendrick Shen, Ananya Kumar, Weihua Hu, Michihiro Yasunaga, Henrik Marklund, Sara Beery, Etienne David, Ian Stavness, Wei Guo, Jure Leskovec, Kate Saenko, Tatsunori Hashimoto, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Later among the works it cites.
Milp modeling of boolean functions by minimum number of inequalities
Aleksei Udovenko · 2021
Later among the works it cites.
On margin maximization in linear and relu networks, 2021
Gal Vardi, Ohad Shamir, and Nathan Srebro · 2021
Later among the works it cites.
Chiyuan Zhang, Maithra Raghu, Jon M. Kleinberg, and Samy Bengio · 2021
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks, 2022
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2022
Closest in time.
An initial alignment between neural network and target is needed for gradient descent to learn, 2022
Emmanuel Abbe, Elisabetta Cornacchia, Jan Hązła, and Christopher Marquis · 2022
Closest in time.
A fine-grained analysis on distribution shift
Olivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi, Ira Ktena, Krishnamurthy Dj Dvijotham, and Ali Taylan Cemgil · 2022
Closest in time.