Fetching the paper…
Reading the bibliography…
This paper provides norm-based generalization bounds for the Transformer architecture that do not depend on the input sequence length.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S. (2019) · 1906
Earlier work this paper cites.
Generalization bounds for convolutional neural networks
Lin, S. and Zhang, J. (2019) · 1910
Earlier work this paper cites.
l ∞ \infty vector contraction for rademacher complexity
Foster, D. J. and Rakhlin, A. (2019) · 1911
Earlier work this paper cites.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
Dudley, R. M. (1967) · 1967
Earlier work this paper cites.
Remarques sur un résultat non publié de b. maurey
Pisier, G. (1981) · 1981
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes
Ledoux, M. and Talagrand, M. (1991) · 1991
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Zhang, T. (2002) · 2002
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Kakade, S. M., Sridharan, K., and Tewari, A. (2008) · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020) · 2010
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C. (2022) · 2022
Later among the works it cites.
On rademacher complexity-based generalization bounds for deep learning
Truong, L. V. (2022) · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Wei, C., Chen, Y., and Ma, T. (2022) · 2022
Later among the works it cites.
Multistep short-term wind speed forecasting using transformer
Wu, H., Meng, K., Fan, D., Zhang, Z., and Liu, Q. (2022) · 2022
Later among the works it cites.
An analysis of attention via the lens of exchangeability and latent variable models
Zhang, Y., Liu, B., Cai, Q., Wang, L., and Wang, Z. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Golowich, N., Rakhlin, A., and Shamir, O. (2018) · 2018
Cited alongside, same era.
Understanding and improving layer normalization
Xu, J., Sun, X., Zhang, Z., Zhao, G., and Lin, J. (2019) · 2019
Cited alongside, same era.
Generalization and representational limits of graph neural networks
Garg, V., Jegelka, S., and Jaakkola, T. (2020) · 2020
Cited alongside, same era.
Fat-shattering dimension of k k -fold maxima
Kontorovich, A. and Attias, I. (2021) · 2021
Cited alongside, same era.
What can a single attention layer learn? a study through the random features lens
Fu, H., Guo, T., Bai, Y., and Mei, S. (2023) · 2023
Closest in time.
Gpt-4 passes the bar exam
Katz, D. M., Bommarito, M. J., Gao, S., and Arredondo, P. (2023) · 2023
Closest in time.
Comparison of lstm, transformers, and mlp-mixer neural networks for gaze based human intention prediction
Pettersson, J. and Falkman, P. (2023) · 2023
Closest in time.