Fetching the paper…
Reading the bibliography…
A type of generalization error induced by initialization in deep neural networks
Zhang, Y., Xu, Z.-Q. J., Luo, T., and Ma, Z · 1905
Earlier work this paper cites.
Gradient dynamics of shallow univariate relu networks
Williams, F., Trager, M., Silva, C. T., Panozzo, D., Zorin, D., and Bruna, J · 1906
Earlier work this paper cites.
Improving deep transformer with depth-scaled initialization and merged attention
Zhang, B., Titov, I., and Sennrich, R · 1908
Earlier work this paper cites.
Reflections after refereeing papers for nips
Breiman, L · 1995
Earlier work this paper cites.
Efficient BackProp , pp. 9–50
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K. R · 1998
Earlier work this paper cites.
The algebraic mind: Integrating connectionism and cognitive science
Marcus, G. F · 2003
Earlier work this paper cites.
Polynomial coefficients and distribution of the sum of discrete uniform variables
Caiado, C. and Rathie, P · 2007
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
Learning compositionally through attentive guidance
Hupkes, D., Singh, A., Korrel, K., Kruszewski, G., and Bruni, E · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Rotskoff, G. and Vanden-Eijnden, E · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J. and Spiliopoulos, K · 2019
Earlier work this paper cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
E, W., Ma, C., and Wu, L · 2020
Cited alongside, same era.
Improving transformer optimization through better initialization
Huang, X. S., Perez, F., Ba, J., and Volkovs, M · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Cited alongside, same era.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Csordás, R., Irie, K., and Schmidhuber, J · 2021
Cited alongside, same era.
Phase diagram for two-layer relu neural networks at infinite-width limit
Luo, T., Xu, Z.-Q. J., Ma, Z., and Zhang, Y · 2021
Cited alongside, same era.
Phase diagram of initial condensation for two-layer neural networks
Chen, Z.-A., Li, Y., Luo, T., Zhou, Z., and Xu, Z.-Q. J · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?, 2023
Eldan, R. and Li, Y · 2023
Later among the works it cites.
Break it down: Evidence for structural compositionality in neural networks
Lepori, M. A., Serre, T., and Pavlick, E · 2023
Later among the works it cites.
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Okawa, M., Lubana, E. S., Dick, R. P., and Tanaka, H · 2023
Later among the works it cites.
How capable can a transformer become? a study on synthetic, interpretable tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradinit: Learning to initialize neural networks for stable and efficient training
Zhu, C., Ni, R., Xu, Z., Kong, K., Huang, W. R., and Goldstein, T · 2021
Cited alongside, same era.
Faithful reasoning using large language models
Creswell, A. and Shanahan, M · 2022
Cited alongside, same era.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Creswell, A., Shanahan, M., and Higgins, I · 2022
Cited alongside, same era.
Csordás, R., Irie, K., and Schmidhuber, J · 2022
Cited alongside, same era.
How does gpt obtain its ability? tracing emergent abilities of language models to their sources
Fu, Y., Peng, H., and Khot, T · 2022
Cited alongside, same era.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Cited alongside, same era.
Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems
Smolensky, P., McCoy, R., Fernandez, R., Goldrick, M., and Gao, J · 2022
Cited alongside, same era.
Ramesh, R., Khona, M., Dick, R. P., Tanaka, H., and Lubana, E. S · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2023
Later among the works it cites.
Mimetic initialization of self-attention layers
Trockman, A. and Kolter, J. Z · 2023
Later among the works it cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Wang, L., Li, L., Dai, D., Chen, D., Zhou, H., Meng, F., Zhou, J., and Sun, X · 2023
Later among the works it cites.
Loss spike in training neural networks
Zhang, Z. and Xu, Z.-Q. J · 2023
Later among the works it cites.
Stochastic modified equations and dynamics of dropout algorithm
Zhang, Z., Li, Y., Luo, T., and Xu, Z.-Q. J · 2023
Later among the works it cites.
Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., Lee, J. R., Lee, Y. T., Li, Y., Liu, W., Mendes, C. C. T., Nguyen, A., Price, E., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Wang, X., Ward, R., Wu, Y., Yu, D., Zhang, C., and Zhang, Y · 2024
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., et al · 2024
Later among the works it cites.
Not all tokens are what you need for pretraining
Lin, Z., Gou, Z., Gong, Y., Liu, X., yelong shen, Xu, R., Lin, C., Yang, Y., Jiao, J., Duan, N., and Chen, W · 2024
Later among the works it cites.
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al · 2024
Later among the works it cites.
Implicit regularization of dropout
Zhang, Z. and Xu, Z.-Q. J · 2024
Later among the works it cites.
An overview of condensation phenomenon in deep learning
Xu, Z.-Q. J., Zhang, Y., and Zhou, Z · 2025
Closest in time.
Complexity control facilitates reasoning-based compositional generalization in transformers
Zhang, Z., Lin, P., Wang, Z., Zhang, Y., and Xu, Z.-Q. J · 2025
Closest in time.