Fetching the paper…
Reading the bibliography…
In-context learning (ICL) of large language models has proven to be a surprisingly effective method of learning a new task from only a few demonstrative examples.
An introduction to matrix concentration inequalities
J. A. Tropp · 1935
Earlier work this paper cites.
Approximation of functions of several variables and imbedding theorems , volume 205 of Grundlehren der mathematischen Wissenschaften
S. M. Nikol’skii · 1975
Earlier work this paper cites.
Nets of Grassmann manifold and orthogonal group
S. J. Szarek · 1981
Earlier work this paper cites.
Theory of function spaces
H. Triebel · 1983
Earlier work this paper cites.
Interpolation of Besov spaces
R. A. DeVore and V. A. Popov · 1988
Earlier work this paper cites.
Weak convergence and empirical processes: with applications to statistics
A. W. van der Vaart and J. A. Wellner · 1996
Earlier work this paper cites.
Minimax estimation via wavelet shrinkage
D. L. Donoho and I. M. Johnstone · 1998
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Y. Yang and A. Barron · 1999
Earlier work this paper cites.
Function spaces with dominating mixed smoothness , volume 30 of Lectures in Mathematics
J. Vybíral · 2006
Earlier work this paper cites.
Widths of embeddings in function spaces
J. Vybíral · 2008
Earlier work this paper cites.
Entropy numbers in function spaces with mixed integrability
H. Triebel · 2011
Earlier work this paper cites.
Mathematical foundations of infinite-dimensional statistical models
E. Giné and R. Nickl · 2015
Earlier work this paper cites.
Error bounds for approximations with deep ReLU networks
D. Yarotsky · 2016
Earlier work this paper cites.
Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality
T. Suzuki · 2019
Earlier work this paper cites.
Low-rank bottleneck in multi-head attention models
S. Bhojanapalli, C. Yun, A. S. Rawat, S. J. Reddi, and S. Kumar · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
S. Hayakawa and T. Suzuki · 2020
Cited alongside, same era.
Nonparametric regression using deep neural networks with ReLU activation function
J. Schmidt-Hieber · 2020
Cited alongside, same era.
Provable meta-learning of linear representations
N. Tripuraneni, C. Jin, and M. Jordan · 2020
Cited alongside, same era.
T. Guo, W. Hu, S. Mei, H. Wang, C. Xiong, S. Savarese, and Y. Bai · 2023
Later among the works it cites.
In-context convergence of Transformers
Y. Huang, Y. Cheng, and Y. Liang · 2023
Later among the works it cites.
A. Mahankali, T. B. Hashimoto, and T. Ma · 2023
Later among the works it cites.
Nonlinear meta-learning can guarantee faster rates
D. Meunier, Z. Li, A. Gretton, and S. Kpotufe · 2023
Later among the works it cites.
Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Chen, T. Dao, E. Winsor, Z. Song, A. Rudra, and C. Ré · 2021
Cited alongside, same era.
Few-shot learning via learning the representation, provably
S. Du, W. Hu, S. Kakade, J. Lee, and Q. Lei · 2021
Cited alongside, same era.
Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov space
T. Suzuki and A. Nitanda · 2021
Cited alongside, same era.
What can Transformers learn in-context? A case study of simple function classes
S. Garg, D. Tsipras, P. Liang, and G. Valiant · 2022
Cited alongside, same era.
Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothness
S. Okumoto and T. Suzuki · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
K. Ahn, X. Cheng, H. Daneshmand, and S. Sra · 2023
Cited alongside, same era.
What learning algorithm is in-context learning? Investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Cited alongside, same era.
A. Raventos, M. Paul, F. Chen, and S. Ganguli · 2023
Later among the works it cites.
Approximation and estimation ability of Transformers for sequence-to-sequence functions with infinite dimensional input
S. Takakura and T. Suzuki · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Later among the works it cites.
Trained Transformers learn linear models in-context
R. Zhang, S. Frei, and P. L. Bartlett · 2023
Later among the works it cites.
S. Chen, H. Sheen, T. Wang, and Z. Yang · 2024
Closest in time.
Transformers learn nonlinear features in context: nonconvex mean-field dynamics on the attention landscape
J. Kim and T. Suzuki · 2024
Closest in time.
H. Li, M. Wang, S. Lu, X. Cui, and P.-Y. Chen · 2024
Closest in time.
Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methods
Y. Nishimura and T. Suzuki · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
J. Wu, D. Zou, Z. Chen, V. Braverman, Q. Gu, and P. L. Bartlett · 2024
Closest in time.
R. Zhang, J. Wu, and P. L. Bartlett · 2024
Closest in time.