Fetching the paper…
Reading the bibliography…
Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context.
Topics in propagation of chaos
Sznitman, A.-S · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A · 1993
Earlier work this paper cites.
A center-stable manifold theorem for differential equations in Banach spaces
Gallay, T · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Barron, A · 1994
Earlier work this paper cites.
The variational formulation of the Fokker–Planck equation
Jordan, R., Kinderlehrer, D., and Otto, F · 1998
Earlier work this paper cites.
The geometry of dissipative evolution equations: the porous medium equation
Otto, F · 2001
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Ambrosio, L., Gigli, N., and Savaré, G · 2005
Earlier work this paper cites.
Noise-induced phenomena in slow-fast dynamical systems: a sample-paths approach
Berglund, N. and Gentz, B · 2006
Earlier work this paper cites.
Some geometric calculations on Wasserstein space
Lott, J · 2008
Earlier work this paper cites.
Optimal Transport: Old and New
Villani, C · 2009
Earlier work this paper cites.
Kernels for vector-valued functions: a review
Álvarez, M., Rosasco, L., and Lawrence, N · 2012
Earlier work this paper cites.
Global Stability of Dynamical Systems
Shub, M · 2013
Earlier work this paper cites.
On the rate of convergence in Wasserstein distance of the empirical measure
Fournier, N. and Guillin, A · 2015
Earlier work this paper cites.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
Santambrogio, F · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Tropp, J. A · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Earlier work this paper cites.
Risk bounds for high-dimensional ridge function combinations including neural networks
Klusowski, J. and Barron, A · 2016
Earlier work this paper cites.
Gradient descent can take exponential time to escape saddle points
Du, S. S., Jin, C., Lee, J., Jordan, M. I., Singh, A., and Póczos, B · 2017
Earlier work this paper cites.
No spurious local minima in nonconvex low rank problems: a unified geometric analysis
Ge, R., Jin, C., and Zheng, Y · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Cited alongside, same era.
Density Ratio Estimation in Machine Learning
Sugiyama, M., Suzuki, T., and Kanamori, T · 2018
Cited alongside, same era.
One-dimensional empirical measures, order statistics, and Kantorovich transport distances
Bobkov, S. G. and Ledoux, M · 2019
Cited alongside, same era.
First-order methods almost always avoid saddle points
Lee, J., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M., and Recht, B · 2019
Cited alongside, same era.
Symmetry, saddle points, and global optimization landscape of nonconvex matrix factorization
Li, X., Lu, J., Arora, R., Haupt, J., Liu, H., Wang, Z., and Zhao, T · 2019
Learning time-scales in two-layers neural networks
Berthier, R., Montanari, A., and Zhou, K · 2023
Later among the works it cites.
On learning Gaussian multi-index models with gradient flow
Bietti, A., Bruna, J., and Pillaud-Vivien, L · 2023
Later among the works it cites.
Guo, T., Hu, W., Mei, S., Wang, H., Xiong, C., Savarese, S., and Bai, Y · 2023
Later among the works it cites.
Explaining emergent in-context learning as kernel regression
Han, C., Wang, Z., Zhao, H., and Ji, H · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Global convergence of neuron birth-death dynamics
Rotskoff, G., Jelassi, S., Bruna, J., and Vanden-Eijnden, E · 2019
Cited alongside, same era.
Transformer dissection: an unified understanding for Transformer’s attention via the lens of kernel
Tsai, Y.-H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Cited alongside, same era.
Regularization matters: generalization and optimization of neural nets v.s. their induced kernel
Wei, C., Lee, J., Liu, Q., and Ma, T · 2019
Cited alongside, same era.
A priori estimates of the population risk for two-layer neural networks
Weinan, E., Ma, C., and Wu, L · 2019
Cited alongside, same era.
Transformers are RNNs: fast autoregressive Transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Complexity measures for neural networks with general activation fnctions using path-based norms
Li, Z., Ma, C., and Wu, L · 2020
Cited alongside, same era.
Huang, Y., Cheng, Y., and Liang, Y · 2023
Later among the works it cites.
Lin, L., Bai, Y., and Mei, S · 2023
Later among the works it cites.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Later among the works it cites.
Leveraging the two-timescale regime to demonstrate convergence of neural networks
Marion, P. and Berthier, R · 2023
Later among the works it cites.
Do pretrained Transformers really learn in-context by gradient descent?
Shen, L., Mishra, A., and Khashabi, D · 2023
Later among the works it cites.
Convergence of mean-field Langevin dynamics: Time and space discretization, stochastic gradient, and variance reduction
Suzuki, T., Wu, D., and Nitanda, A · 2023
Later among the works it cites.
JoMA: demystifying multilayer Transformers via joint dynamics of MLP and attention, 2023
Tian, Y., Wang, Y., Zhang, Z., Chen, B., and Du, S · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Gated linear attention Transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Later among the works it cites.
Sequential attention for feature selection
Yasuda, T., Bateni, M., Chen, L., Fahrbach, M., Fu, G., and Mirrokni, V · 2023
Later among the works it cites.
On the global convergence of Wasserstein gradient flow of the Coulomb discrepancy
Boufadène, S. and Vialard, F.-X · 2024
Closest in time.
Chen, S., Sheen, H., Wang, T., and Yang, Z · 2024
Closest in time.
Symmetric mean-field Langevin dynamics for distributional minimax problems
Kim, J., Yamamoto, K., Oko, K., Yang, Z., and Suzuki, T · 2024
Closest in time.
Li, H., Wang, M., Lu, S., Cui, X., and Chen, P.-Y · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L · 2024
Closest in time.
Zhang, R., Wu, J., and Bartlett, P. L · 2024
Closest in time.