Fetching the paper…
Reading the bibliography…
Modern deep neural networks display striking examples of rich internal computational structure.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
Inference from iterative simulation using multiple sequences
Gelman, A. and Rubin, D. B · 1992
Earlier work this paper cites.
Evaluating the accuracy of sampling-based approaches to the calculations of posterior moments
Geweke, J · 1992
Earlier work this paper cites.
Essential dynamics of proteins
Amadei, A., Linssen, A. B., and Berendsen, H. J · 1993
Earlier work this paper cites.
On the problem of applying AIC to determine the structure of a layered feedforward neural network
Hagiwara, K., Toda, N., and Usui, S · 1993
Earlier work this paper cites.
Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering
Strogatz, S. H · 1994
Earlier work this paper cites.
Similarities between principal components of protein dynamics and random diffusion
Hess, B · 2000
Earlier work this paper cites.
Bayesian Data Analysis
Gelman, A., Carlin, J. B., Stern, H. S., and Rubin, D. B · 2003
Earlier work this paper cites.
Semantic Cognition: A Parallel Distributed Processing Approach
Rogers, T. T. and McClelland, J. L · 2004
Earlier work this paper cites.
Optical imaging of neuronal populations during decision-making
Briggman, K. L., Abarbanel, H. D., and Kristan Jr, W · 2005
Earlier work this paper cites.
Essential dynamics: a tool for efficient trajectory compression and management
Meyer, T., Ferrer-Costa, C., Pérez, A., Rueda, M., Bidon-Chanal, A., Luque, F. J., Laughton, C. A., and Orozco, M · 2006
Earlier work this paper cites.
Almost all learning machines are singular
Watanabe, S · 2007
Earlier work this paper cites.
Normal modes and essential dynamics
Hayward, S. and De Groot, B. L · 2008
Earlier work this paper cites.
Karhunen-Loeve Expansions and Their Applications
Wang, L · 2008
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Watanabe, S · 2009
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Brain-wide neuronal dynamics during motor adaptation in zebrafish
Ahrens, M. B., Li, J. M., Orger, M. B., Robson, D. N., Schier, A. F., Engert, F., and Portugues, R · 2012
Earlier work this paper cites.
Dimensionality reduction for large-scale neural recordings
Cunningham, J. P. and Yu, B. M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N · 2017
Earlier work this paper cites.
PCA of high dimensional random walks with comparison to neural network training
Antognini, J. and Sohl-Dickstein, J · 2018
Cited alongside, same era.
Mathematical Theory of Bayesian Statistics
Watanabe, S · 2018
Cited alongside, same era.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Cited alongside, same era.
SGD on neural networks learns functions of increasing complexity
Kalimeris, D., Kaplun, G., Nakkiran, P., Edelman, B., Yang, T., Barak, B., and Zhang, H · 2019
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2019
Cited alongside, same era.
Thread: Circuits
Cammarata, N., Olah, C., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
SGD learning on neural networks: Leap complexity and saddle-to-saddle dynamics
Abbe, E., Adserà, E. B., and Misiakiewicz, T · 2023
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models, 2023
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Later among the works it cites.
Dynamical versus Bayesian phase transitions in a toy model of superposition, 2023
Chen, Z., Lau, E., Mendel, J., Wei, S., and Murfet, D · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventós, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The implicit bias of depth: How incremental learning drives generalization
Gissin, D., Shalev-Shwartz, S., and Daniely, A · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
Revisiting the Gelman–Rubin diagnostic, 2020
Vats, D. and Knudson, C · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Jacot, A., Ged, F., Şimşek, B., Hongler, C., and Gabriel, F · 2021
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F · 2022
Cited alongside, same era.
Phantom oscillations in principal component analysis
Shinn, M · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2024
Later among the works it cites.
Toward understanding in-context vs. in-weight learning, 2024
Chan, B., Chen, X., György, A., and Schuurmans, D · 2024
Later among the works it cites.
The evolution of statistical induction heads: In-context learning markov chains
Edelman, E., Tsilivis, N., Edelman, B. L., Malach, E., and Goel, S · 2024
Later among the works it cites.
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
He, T., Doshi, D., Das, A., and Gromov, A · 2024
Later among the works it cites.
The developmental landscape of in-context learning, 2024
Hoogland, J., Wang, G., Farrugia-Roberts, M., Carroll, L., Wei, S., and Murfet, D · 2024
Later among the works it cites.
The training process of many deep networks explores the same low-dimensional manifold
Mao, J., Griniasty, I., Teoh, H. K., Ramesh, R., Yang, R., Transtrum, M. K., Sethna, J. P., and Chaudhari, P · 2024
Later among the works it cites.
Nguyen, A. and Reddy, G · 2024
Later among the works it cites.
In-context learning through the Bayesian prism
Panwar, M., Ahuja, K., and Goyal, N · 2024
Later among the works it cites.
Competition dynamics shape algorithmic phases of in-context learning, 2024
Park, C. F., Lubana, E. S., Pres, I., and Tanaka, H · 2024
Later among the works it cites.
The transient nature of emergent in-context learning in transformers
Singh, A., Chan, S., Moskovitz, T., Grant, E., Saxe, A., and Hill, F · 2024
Later among the works it cites.
Wang, G., Hoogland, J., van Wingerden, S., Furman, Z., and Murfet, D · 2024
Later among the works it cites.
Hyperparameter tuning for local learning coefficient estimation
Xu, A. K · 2024
Later among the works it cites.
The local learning coefficient: A singularity-aware complexity measure
Lau, E., Furman, Z., Wang, G., Murfet, D., and Wei, S · 2025
Closest in time.