Fetching the paper…
Reading the bibliography…
Classes of target functions containing a large number of approximately orthogonal elements are known to be hard to learn by the Statistical Query algorithms.
Boas, R.: A general moment problem. American Journal of Mathematics 63
1941
Earlier work this paper cites.
Bellman, R.: Almost orthogonal series. Bulletin of the American Mathematical Society 50
1944
Earlier work this paper cites.
Sinai, Y.G.: Introduction to Ergodic Theory. Mathematical notes; 18. Princeton University Press, Princeton (1977)
1977
Earlier work this paper cites.
Kearns, M.J.: Efficient noise-tolerant learning from statistical queries. In: Kosaraju, S.R., Johnson, D.S., Aggarwal, A. (eds.) Proceedings of the Twenty-Fifth Annual ACM Symposium on Theory of Computing, May 16-18, 1993, San Diego, CA, USA, pp. 392–401. ACM, ??? (1993). https://doi.org/10.1145/167088.167200 . https://doi.org/10.1145/167088.167200
1993
Earlier work this paper cites.
Blum, A., Furst, M.L., Jackson, J.C., Kearns, M.J., Mansour, Y., Rudich, S.: Weakly learning DNF and characterizing statistical query learning using fourier analysis. In: Leighton, F.T., Goodrich, M.T. (eds.) Proceedings of the Twenty-Sixth Annual ACM Symposium on Theory of Computing, 23-25 May 1994, Montréal, Québec, Canada, pp. 253–262. ACM, ??? (1994). https://doi.org/10.1145/195058.195147 . https://doi.org/10.1145/195058.195147
1994
Earlier work this paper cites.
Hochreiter, S., Schmidhuber, J.: Long Short-Term Memory. Neural Computation 9
1997
Earlier work this paper cites.
Barthe, F.: Optimal young’s inequality and its converse: a simple proof. Geometric & Functional Analysis GAFA 8
1998
Earlier work this paper cites.
Yang, K.: On learning correlated boolean functions using statistical query. In: Proceedings of the Algorithmic Learning Theory, 12th International Conference, ALT 2001, Washington, DC, USA, pp. 59–76 (2001)
2001
Earlier work this paper cites.
Klivans, A.R., Sherstov, A.A.: Cryptographic hardness for learning intersections of halfspaces. Journal of Computer and System Sciences 75
2006
Earlier work this paper cites.
Klivans, A.R., Sherstov, A.A.: Unconditional lower bounds for learning intersections of halfspaces. Mach. Learn. 69
2007
Earlier work this paper cites.
Kuipers, L., Niederreiter, H.: Uniform Distribution of Sequences. Dover Books on Mathematics. Dover Publications, ??? (2012). https://books.google.kz/books?id=mnY8LpyXHM0C
2012
Cited alongside, same era.
Livni, R., Shalev-Shwartz, S., Shamir, O.: On the computational efficiency of training neural networks. In: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K.Q. (eds.) Advances in Neural Information Processing Systems, vol. 27. Curran Associates, Inc., ??? (2014)
2014
Cited alongside, same era.
He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition (2015)
2015
Cited alongside, same era.
Alon, N., Shikhelman, C.: Many t copies in h-free graphs. Journal of Combinatorial Theory, Series B 121
2016
Cited alongside, same era.
Hardt, M., Ma, T.: Identity matters in deep learning. In: International Conference on Learning Representations (2017). https://openreview.net/forum?id=ryxB0Rtxx
Takhanov, R.: Reducing the dimensionality of data using tempered distributions. Digital Signal Processing 133
2022
Later among the works it cites.
Liu, Z., Yu, L.-W., Duan, L.-M., Deng, D.-L.: Presence and absence of barren plateaus in tensor-network based machine learning. Phys. Rev. Lett. 129
2022
Later among the works it cites.
Power, A., Burda, Y., Edwards, H., Babuschkin, I., Misra, V.: Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets (2022)
2022
Later among the works it cites.
Wenger, E., Chen, M., Charton, F., Lauter, K.: SALSA: Attacking Lattice Cryptography with Transformers. Cryptology ePrint Archive, Paper 2022/935 (2022). https://eprint.iacr.org/2022/935
2022
Later among the works it cites.
Takhanov, R., Abylkairov, Y.S., Tezekbayev, M.: Autoencoders for a manifold learning problem with a jacobian rank constraint. Pattern Recognition 143
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Feldman, V., Guzmán, C., Vempala, S.S.: Statistical query algorithms for mean vector estimation and stochastic convex optimization. In: Klein, P.N. (ed.) Proceedings of the Twenty-Eighth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2017, Barcelona, Spain, Hotel Porta Fira, January 16-19, pp. 1265–1277. SIAM, ??? (2017). https://doi.org/10.1137/1.9781611974782.82 . https://doi.org/10.1137/1.9781611974782.82
2017
Cited alongside, same era.
Shalev-Shwartz, S., Shamir, O., Shammah, S.: Failures of gradient-based deep learning. In: Precup, D., Teh, Y.W. (eds.) Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017. Proceedings of Machine Learning Research, vol. 70, pp. 3067–3075. PMLR, ??? (2017). http://proceedings.mlr.press/v70/shalev-shwartz17a.html
2017
Cited alongside, same era.
Assylbekov, Z., Takhanov, R.: Reusing weights in subword-aware neural language models. In: Walker, M., Ji, H., Stent, A. (eds.) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1413–1423. Association for Computational Linguistics, New Orleans, Louisiana (2018). https://doi.org/10.18653/v1/N18-1128 . https://aclanthology.org/N18-1128
2018
Cited alongside, same era.
McClean, J.R., Boixo, S., Smelyanskiy, V.N., Babbush, R., Neven, H.: Barren plateaus in quantum neural network training landscapes. Nature Communications 9
2018
Cited alongside, same era.
Shamir, O.: Distribution-specific hardness of learning neural networks. J. Mach. Learn. Res. 19
2018
Cited alongside, same era.
2023
Closest in time.
Gromov, A.: Grokking modular arithmetic (2023)
2023
Closest in time.
Liu, Z., Michaud, E.J., Tegmark, M.: Omnigrok: Grokking beyond algorithmic data. In: The Eleventh International Conference on Learning Representations (2023). https://openreview.net/forum?id=zDiHoIWa0q1
2023
Closest in time.
Takhanov, R.: Multi-layer random features and the approximation power of neural networks. In: Evans, R.J., Shpitser, I. (eds.) Proceedings of the Fourtieth Conference on Uncertainty in Artificial Intelligence. Proceedings of Machine Learning Research, vol. 248, pp. 33–44. PMLR, ??? (2024)
2024
Closest in time.
Takhanov, R., Tezekbayev, M., Pak, A., Bolatov, A., Kadyrsizova, Z., Assylbekov, Z.: Intractability of learning the discrete logarithm with gradient-based methods. In: Yanıkoğlu, B., Buntine, W. (eds.) Proceedings of the 15th Asian Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 222, pp. 1321–1336. PMLR, ??? (2024). https://proceedings.mlr.press/v222/takhanov24a.html
2024
Closest in time.