Fetching the paper…
Reading the bibliography…
In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs.
Learning polynomials with neural networks
Andoni, A., Panigrahy, R., Valiant, G., and Zhang, L. (2014) · 1916
Earlier work this paper cites.
Perceptrons: an introduction to computational geometry
Minsky, M. and Papert, S. (1969) · 1969
Earlier work this paper cites.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Blum, A., Furst, M., Jackson, J., Kearns, M., Mansour, Y., and Rudich, S. (1994) · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Kearns, M. (1998) · 1998
Earlier work this paper cites.
Random forests
Breiman, L. (2001) · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Mutual learning in a tree parity machine and its application to cryptography
Rosen-Zvi, M., Klein, E., Kanter, I., and Kinzel, W. (2002) · 2002
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Blum, A., Kalai, A., and Wasserman, H. (2003) · 2003
Earlier work this paper cites.
Evolvability from learning algorithms
Feldman, V. (2008) · 2008
Earlier work this paper cites.
Fast cryptographic primitives and circular-secure encryption based on hard learning problems
Applebaum, B., Cash, D., Peikert, C., and Sahai, A. (2009) · 2009
Earlier work this paper cites.
Public-key cryptography from different assumptions
Applebaum, B., Barak, B., and Wigderson, A. (2010) · 2010
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020) · 2010
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al. (2020) · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. (2011) · 2011
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T. (2013) · 2013
Earlier work this paper cites.
Analysis of Boolean functions
O’Donnell, R. (2014) · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Statistical query algorithms for mean vector estimation and stochastic convex optimization
Feldman, V., Guzman, C., and Vempala, S. (2017) · 2017
Earlier work this paper cites.
Time-space hardness of learning sparse parities
Kol, G., Raz, R., and Tal, A. (2017) · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shalev-Shwartz, S., Shamir, O., and Shammah, S. (2017) · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2018) · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M. (2018) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Abbe, E., Adsera, E. B., and Misiakiewicz, T. (2022) · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Ba, J., Erdogdu, M. A., Suzuki, T., Wang, Z., Wu, D., and Yang, G. (2022) · 2022
Later among the works it cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B., Edelman, B., Goel, S., Kakade, S., Malach, E., and Zhang, C. (2022) · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Bietti, A., Bruna, J., Sanford, C., and Song, M. J. (2022) · 2022
Later among the works it cites.
Deep neural networks and tabular data: A survey
Borisov, V., Leemann, T., Seßler, K., Haug, J., Pawelczyk, M., and Kasneci, G. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T. (2019) · 2019
Cited alongside, same era.
Learning polynomials in few relevant dimensions
Chen, S. and Meka, R. (2020) · 2020
Cited alongside, same era.
Learning parities with neural networks
Daniely, A. and Malach, E. (2020) · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hahn, M. (2020) · 2020
Cited alongside, same era.
Approximate is good enough: Probabilistic variants of dimensional and margin complexity
Kamath, P., Montasser, O., and Srebro, N. (2020) · 2020
Cited alongside, same era.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U. (2021) · 2021
Cited alongside, same era.
Neural networks can learn representations with gradient descent
Damian, A., Lee, J., and Soltanolkotabi, M. (2022) · 2022
Later among the works it cites.
Random feature amplification: Feature learning and generalization in neural networks
Frei, S., Chatterji, N. S., and Bartlett, P. L. (2022) · 2022
Later among the works it cites.
Why do tree-based models still outperform deep learning on typical tabular data?
Grinsztajn, L., Oyallon, E., and Varoquaux, G. (2022) · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Later among the works it cites.
Omnigrok: Grokking beyond algorithmic data
Liu, Z., Michaud, E. J., and Tegmark, M. (2022) · 2022
Later among the works it cites.
Sparse tree-based initialization for neural networks
Lutz, P., Arnould, L., Boyer, C., and Scornet, E. (2022) · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V. (2022) · 2022
Later among the works it cites.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Telgarsky, M. (2022) · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. (2022) · 2022
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei, S., Wei, J., and Liang, Y. (2022) · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Abbe, E., Boix-Adsera, E., and Misiakiewicz, T. (2023) · 2023
Closest in time.
Scaling mlps: A tale of inductive bias
Bachmann, G., Anagnostidis, S., and Hofmann, T. (2023) · 2023
Closest in time.
Damian, A., Nichani, E., Ge, R., and Lee, J. D. (2023) · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
Merrill, W., Tsilivis, N., and Shukla, A. (2023) · 2023
Closest in time.
The quantization model of neural scaling
Michaud, E. J., Liu, Z., Girit, U., and Tegmark, M. (2023) · 2023
Closest in time.