Fetching the paper…
Reading the bibliography…
The $k$-sparse parity problem is a classical problem in computational complexity and algorithmic theory, serving as a key benchmark for understanding computational classes.
Mean-field langevin dynamics and energy landscape of neural networks
Hu, K · 1905
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The rprop algorithm
Riedmiller, M · 1993
Earlier work this paper cites.
The intractability of computing the minimum distance of a code
Vardy, A · 1997
Earlier work this paper cites.
Limit on the speed of quantum computation in determining parity
Farhi, E · 1998
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Kearns, M · 1998
Earlier work this paper cites.
The parametrized complexity of some fundamental problems in coding theory
Downey, R. G · 1999
Earlier work this paper cites.
The calculus of finite differences
Milne-Thomson, L. M · 2000
Earlier work this paper cites.
On using extended statistical queries to avoid membership queries
Bshouty, N. H · 2002
Earlier work this paper cites.
Hardness of approximating the minimum distance of a linear code
Dumer, I · 2003
Earlier work this paper cites.
On-line algorithms in machine learning
Blum, A · 2005
Earlier work this paper cites.
Toward attribute efficient learning of decision lists and parities
Klivans, A. R · 2006
Earlier work this paper cites.
A tight lower bound for parity in noisy communication networks
Dutta, C · 2008
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L · 2018
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S · 2018
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Wei, C · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L · 2020
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B · 2022
Later among the works it cites.
Random feature amplification: Feature learning and generalization in neural networks
Frei, S · 2022
Later among the works it cites.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Telgarsky, M · 2022
Later among the works it cites.
Glasgow, M · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning parities with neural networks
Daniely, A · 2020
Cited alongside, same era.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ji, Z · 2020
Cited alongside, same era.
Understanding the generalization of adam in learning neural networks with proper regularization
Zou, D · 2021
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Abbe, E · 2022
Cited alongside, same era.
On the non-universality of deep learning: quantifying the cost of symmetry
Abbe, E · 2022
Cited alongside, same era.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Abbe, E
Cited in the paper.
Liu, H · 2023
Later among the works it cites.
Benign overfitting in two-layer relu convolutional neural networks for xor data
Meng, X · 2023
Later among the works it cites.
Feature learning via mean-field langevin dynamics: classifying sparse parities and beyond
Suzuki, T · 2023
Later among the works it cites.
Benign overfitting and grokking in relu networks for xor cluster data
Xu, Z · 2023
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X · 2024
Closest in time.
Pareto frontiers in deep feature learning: Data, compute, width, and luck
Edelman, B · 2024
Closest in time.
Improved statistical and computational complexity of the mean-field langevin dynamics under structured data
Nitanda, A · 2024
Closest in time.