Fetching the paper…
Reading the bibliography…
Recent advances in deep learning have mainly relied on Transformers due to their data dependency and ability to learn at scale.
Object categorization
Pinz, A. et al · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
State space modeling of time series
Aoki, M · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Sparse convolutional neural networks
Liu, B., Wang, M., Foroosh, H., Tappen, M., and Pensky, M · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Chollet, F · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Martin, E. and Cundy, C · 2018
Earlier work this paper cites.
Mmdetection: Open mmlab detection toolbox and benchmark
Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., et al · 2019
Earlier work this paper cites.
Depthwise convolution is all you need for learning multiple visual domains
Guo, Y., Li, Y., Wang, L., and Rosing, T · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., and Torralba, A · 2019
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Earlier work this paper cites.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Cited alongside, same era.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T · 2020
Cited alongside, same era.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Temporal fusion transformers for interpretable multi-horizon time series forecasting
Lim, B., Arık, S. Ö., Loeff, N., and Pfister, T · 2021
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
A time series is worth 64 words: Long-term forecasting with transformers
Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Ré, C · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Later among the works it cites.
Convolutional state space models for long-range spatiotemporal modeling
Smith, J. T., Mello, S. D., Kautz, J., Linderman, S., and Byeon, W · 2023
Later among the works it cites.
Brain encoding models based on multimodal transformers can transfer across language and vision
Tang, J., Du, M., Vo, V., LAL, V., and Huth, A · 2023
Later among the works it cites.
Patches are all you need?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Cited alongside, same era.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Wu, H., Xu, J., Wang, J., and Long, M · 2021
Cited alongside, same era.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W · 2021
Cited alongside, same era.
Monarch: Expressive structured matrices for efficient and accurate training
Dao, T., Chen, B., Sohoni, N. S., Desai, A., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
M5 accuracy competition: Results, findings, and conclusions
Makridakis, S., Spiliotis, E., and Assimakopoulos, V · 2022
Cited alongside, same era.
S4nd: Modeling images and videos as multidimensional signals with state spaces
Nguyen, E., Goel, K., Gu, A., Downs, G., Shah, P., Dao, T., Baccus, S., and Ré, C · 2022
Cited alongside, same era.
Trockman, A. and Kolter, J. Z · 2023
Later among the works it cites.
Effectively modeling time series with simple discrete state spaces
Zhang, M., Saab, K. K., Poli, M., Dao, T., Goel, K., and Re, C · 2023
Later among the works it cites.
Graph mamba: Towards learning on graphs with state space models
Behrouz, A. and Hashemi, F · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., et al · 2024
Closest in time.
Zigma: Zigzag mamba diffusion model
Hu, V. T., Baumann, S. A., Gui, M., Grebenkova, O., Ma, P., Fischer, J., and Ommer, B · 2024
Closest in time.
Localmamba: Visual state space model with windowed selective scan
Huang, T., Pei, X., You, S., Wang, F., Qian, C., and Xu, C · 2024
Closest in time.
Ilbert, R., Odonnat, A., Feofanov, V., Virmaux, A., Paolo, G., Palpanas, T., and Redko, I · 2024
Closest in time.
Orchid: Flexible and data-dependent convolution for sequence modeling
Karami, M. and Ghodsi, A · 2024
Closest in time.
Mamba-nd: Selective state space modeling for multi-dimensional data
Li, S., Singh, H., and Grover, A · 2024
Closest in time.
Vmamba: Visual state space model
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., and Liu, Y · 2024
Closest in time.
Denseformer: Enhancing information flow in transformers via depth weighted averaging
Pagliardini, M., Mohtashami, A., Fleuret, F., and Jaggi, M · 2024
Closest in time.
Simba: Simplified mamba-based architecture for vision and multivariate time series, 2024
Patro, B. N. and Agneeswaran, V. S · 2024
Closest in time.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Schiff, Y., Kao, C.-H., Gokaslan, A., Dao, T., Gu, A., and Kuleshov, V · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., and Wang, X · 2024
Closest in time.