Fetching the paper…
Reading the bibliography…
Recent advances in deep learning have relied heavily on the use of large Transformers due to their ability to learn at scale.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, C. Hillier, and T. P. Lillicrap · 1911
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
K. Fukushima and S. Miyake · 1982
Earlier work this paper cites.
Linear system theory and design
C.-T. Chen · 1984
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Signals and systems , volume 2
A. V. Oppenheim, A. S. Willsky, S. H. Nawab, and J.-J. Ding · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Butterfly factorization
Y. Li, H. Yang, E. R. Martin, K. L. Ho, and L. Ying · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Fast convolution and filtering
I. W. Selesnick and C. S. Burrus · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Learning fast algorithms for linear transforms using butterfly factorizations
T. Dao, A. Gu, M. Eichhorn, A. Rudra, and C. Ré · 2019
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
D. Hendrycks, N. Mu, E. D. Cubuk, B. Zoph, J. Gilmer, and B. Lakshminarayanan · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo · 2019
Cited alongside, same era.
Frequency bias in neural networks for input of non-uniform density
R. Basri, M. Galun, A. Geifman, D. Jacobs, Y. Kasten, and S. Kritchman · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Later among the works it cites.
RWKV-LM, 8 2021
B. Peng · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
I. Schlag, K. Irie, and J. Schmidhuber · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Later among the works it cites.
S. Zhai, W. Talbott, N. Srivastava, C. Huang, H. Goh, R. Zhang, and J. Susskind · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Cited alongside, same era.
Reformer: The efficient transformer
N. Kitaev, Ł. Kaiser, and A. Levskaya · 2020
Cited alongside, same era.
Fourier neural operator for parametric partial differential equations
Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar · 2020
Cited alongside, same era.
Dissecting neural odes
S. Massaroli, M. Poli, J. Park, A. Yamashita, and H. Asama · 2020
Cited alongside, same era.
Implicit neural representations with periodic activation functions
V. Sitzmann, J. N. Martel, A. W. Bergman, D. B. Lindell, and G. Wetzstein · 2020
Cited alongside, same era.
Later among the works it cites.
Ask me anything: A simple strategy for prompting language models
S. Arora, A. Narayan, M. F. Chen, L. J. Orr, N. Guha, K. Bhatia, I. Chami, F. Sala, and C. Ré · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. Liang, and G. Valiant · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
What makes convolutional models great on long sequence modeling?
Y. Li, T. Cai, Y. Zhang, D. Chen, and D. Dey · 2022
Later among the works it cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
H. Mehta, A. Gupta, A. Cutkosky, and B. Neyshabur · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals using state spaces
E. Nguyen, K. Goel, A. Gu, G. W. Downs, P. Shah, T. Dao, S. A. Baccus, and C. Ré · 2022
Later among the works it cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
Maxvit: Multi-axis vision transformer
Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y. Li · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Y. Zhang, A. Backurs, S. Bubeck, R. Eldan, S. Gunasekar, and T. Wagner · 2022
Later among the works it cites.