High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J · 2019
Later among the works it cites.
Learning one-hidden-layer ReLU networks via gradient descent
Zhang, X · 2019
Later among the works it cites.
Probabilistic symmetries and invariant neural networks
Bloem-Reddy, B · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T · 2020
Later among the works it cites.
An image is worth 16 × \times 16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Later among the works it cites.
Infinite attention: NNGP and NTK for deep attention networks
Hron, J · 2020
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T · 2020
Later among the works it cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Original
Bronstein, M. M · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L · 2021
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Original
Edelman, B. L · 2021
Later among the works it cites.
Provably strict generalisation benefit for invariance in kernel methods
Elesedy, B · 2021
Later among the works it cites.
Lietransformer: Equivariant self-attention for Lie groups
Hutchinson, M. J · 2021
Later among the works it cites.
Highly accurate protein structure prediction with AlphaFold
Jumper, J · 2021
Later among the works it cites.
Self-attention between datapoints: Going beyond individual input-output pairs in deep learning
Kossen, J · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A · 2021
Later among the works it cites.
Improved generalization bounds of group invariant/equivariant deep networks via quotient feature spaces
Sannai, A · 2021
Later among the works it cites.
E ( n ) {E}(n) equivariant normalizing flows
Original
Satorras, V. G · 2021
Later among the works it cites.
Statistically meaningful approximation: A case study on approximating Turing machines with transformers
Original
Wei, C · 2021
Later among the works it cites.
An explanation of in-context learning as implicit Bayesian inference
Original
Xie, S. M · 2021
Later among the works it cites.
Tensor programs IIb: Architectural universality of neural tangent kernel training dynamics
Yang, G · 2021
Later among the works it cites.
Understanding the generalization benefit of model invariance from a data perspective
Zhu, S · 2021
Later among the works it cites.
What can transformers learn in-context? A case study of simple function classes
Original
Garg, S · 2022
Closest in time.
Geometrically equivariant graph neural networks: A survey
Original
Han, J · 2022
Closest in time.
Masked autoencoders are scalable vision learners
He, K · 2022
Closest in time.
A kernel-based view of language model fine-tuning
Original
Malladi, S · 2022
Closest in time.