Fetching the paper…
Reading the bibliography…
In recent years, Vision Transformers have attracted increasing interest from computer vision researchers.
A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K. 1980 · 1980
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y.; Boser, B.; Denker, J. S.; Henderson, D.; Howard, R. E.; Hubbard, W.; and Jackel, L. D. 1989 · 1989
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Cohen, N.; Sharir, O.; and Shashua, A. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; and Schmid, C. 2021 · 2021
Earlier work this paper cites.
Coatnet: Marrying convolution and attention for all data sizes
Dai, Z.; Liu, H.; Le, Q. V.; and Tan, M. 2021 · 2021
Earlier work this paper cites.
Convit: Improving vision transformers with soft convolutional inductive biases
d’Ascoli, S.; Touvron, H.; Leavitt, M. L.; Morcos, A. S.; Biroli, G.; and Sagun, L. 2021 · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A.; Goel, K.; and Ré, C. 2021 · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Gu, A.; Johnson, I.; Goel, K.; Saab, K.; Dao, T.; Rudra, A.; and Ré, C. 2021 · 2021
Earlier work this paper cites.
Vision transformer for small-size datasets
Lee, S. H.; Lee, S.; and Song, B. C. 2021 · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Ckconv: Continuous kernel convolution for sequential data
Romero, D. W.; Kuzina, A.; Bekkers, E. J.; Tomczak, J. M.; and Hoogendoorn, M. 2021 · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Cited alongside, same era.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021 · 2021
Cited alongside, same era.
What Makes Convolutional Models Great on Long Sequence Modeling?
Li, Y.; Cai, T.; Zhang, Y.; Chen, D.; and Dey, D. 2022 · 2022
Later among the works it cites.
Mega: moving average equipped gated attention
Ma, X.; Zhou, C.; Kong, X.; He, J.; Gui, L.; Neubig, G.; May, J.; and Zettlemoyer, L. 2022 · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
Mehta, H.; Gupta, A.; Cutkosky, A.; and Neyshabur, B. 2022 · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals with state spaces
Nguyen, E.; Goel, K.; Gu, A.; Downs, G.; Shah, P.; Dao, T.; Baccus, S.; and Ré, C. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Xu, Y.; Zhang, Q.; Zhang, J.; and Tao, D. 2021 · 2021
Cited alongside, same era.
Tensor decompositions: computations, applications, and challenges
Bi, Y.; Lu, Y.; Long, Z.; Zhu, C.; and Liu, Y. 2022 · 2022
Cited alongside, same era.
Conditional detr v2: Efficient detection transformer with box queries
Chen, X.; Wei, F.; Zeng, G.; and Wang, J. 2022 · 2022
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models
Dao, T.; Fu, D. Y.; Saab, K. K.; Thomas, A. W.; Rudra, A.; and Ré, C. 2022 · 2022
Cited alongside, same era.
Decision S4: Efficient Sequence-Based RL via State Spaces Layers
David, S. B.; Zimerman, I.; Nachmani, E.; and Wolf, L. 2022 · 2022
Cited alongside, same era.
Cmt: Convolutional neural networks meet vision transformers
Guo, J.; Han, K.; Wu, H.; Tang, Y.; Chen, X.; Wang, Y.; and Xu, C. 2022 · 2022
Cited alongside, same era.
Diagonal state spaces are as effective as structured state spaces
Gupta, A.; Gu, A.; and Berant, J. 2022 · 2022
Cited alongside, same era.
Wang, J.; Yan, J. N.; Gu, A.; and Rush, A. M. 2022 · 2022
Later among the works it cites.
2-D SSM: A General Spatial Layer for Visual Transformers
Baron, E.; Zimerman, I.; and Wolf, L. 2023 · 2023
Closest in time.
Simple hardware-efficient long convolutions for sequence modeling
Fu, D. Y.; Epstein, E. L.; Nguyen, E.; Thomas, A. W.; Zhang, M.; Dao, T.; Rudra, A.; and Ré, C. 2023 · 2023
Closest in time.
Structured state space models for in-context reinforcement learning
Lu, C.; Schroecker, Y.; Gu, A.; Parisotto, E.; Foerster, J.; Singh, S.; and Behbahani, F. 2023 · 2023
Closest in time.
Focus Your Attention (with Adaptive IIR Filters)
Lutati, S.; Zimerman, I.; and Wolf, L. 2023 · 2023
Closest in time.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Nguyen, E.; Poli, M.; Faizi, M.; Thomas, A.; Birch-Sykes, C.; Wornow, M.; Patel, A.; Rabideau, C.; Massaroli, S.; Bengio, Y.; et al. 2023 · 2023
Closest in time.
RWKV: Reinventing RNNs for the Transformer Era
Peng, B.; Alcaide, E.; Anthony, Q.; Albalak, A.; Arcadinho, S.; Cao, H.; Cheng, X.; Chung, M.; Grella, M.; GV, K. K.; et al. 2023 · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Poli, M.; Massaroli, S.; Nguyen, E.; Fu, D. Y.; Dao, T.; Baccus, S.; Bengio, Y.; Ermon, S.; and Ré, C. 2023 · 2023
Closest in time.
Diagonal state space augmented transformers for speech recognition
Saon, G.; Gupta, A.; and Cui, X. 2023 · 2023
Closest in time.