Fetching the paper…
Reading the bibliography…
Transformer architectures rely on position encodings to model the spatial structure of input data.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
Representation theory: a first course , volume 129
Fulton, W. and Harris, J · 2013
Earlier work this paper cites.
Esteves, C., Allen-Blanchette, C., Zhou, X., and Daniilidis, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Harmonic networks: Deep translation and rotation equivariance
Worrall, D. E., Garbin, S. J., Turmukhambetov, D., and Brostow, G. J · 2017
Earlier work this paper cites.
Learning so (3) equivariant representations with spherical cnns
Esteves, C., Allen-Blanchette, C., Makadia, A., and Daniilidis, K · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Pytorch lightning
Falcon, W. A · 2019
Earlier work this paper cites.
Reparameterizing distributions on lie groups
Falorsi, L., de Haan, P., Davidson, T. R., and Forré, P · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 2019
Earlier work this paper cites.
Equivariant transformer networks
Tai, K. S., Bailis, P., and Valiant, G · 2019
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Construction of a machine learning dataset through collaboration: the rsna 2019 brain ct hemorrhage challenge
Flanders, A. E., Prevedello, L. M., Shih, G., Halabi, S. S., Kalpathy-Cramer, J., Ball, R., Mongan, J. T., Stein, A., Kitamura, F. C., Lungren, M. P., et al · 2020
Cited alongside, same era.
Se (3)-transformers: 3d roto-translation equivariant attention networks
Fuchs, F., Worrall, D., Fischer, V., and Welling, M · 2020
Cited alongside, same era.
Differential geometry and lie groups , volume 12
Gallier, J. Q. and Quaintance, J · 2020
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Visionllama: A unified llama interface for vision tasks
Chu, X., Su, J., Zhang, B., and Shen, C · 2024
Closest in time.
Longrope: Extending llm context window beyond 2 million tokens
Ding, Y., Zhang, L. L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
Improving vision transformers to learn small-size dataset from scratch
Lee, S., Lee, S., and Song, B. C · 2022
Cited alongside, same era.
Equiformer: Equivariant graph attention transformer for 3d atomistic graphs
Liao, Y.-L. and Smidt, T · 2022
Cited alongside, same era.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al · 2022
Cited alongside, same era.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Cited alongside, same era.
Deit iii: Revenge of the vit
Touvron, H., Cord, M., and Jégou, H · 2022
Cited alongside, same era.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Cited alongside, same era.
Closest in time.
Contextual position encoding: Learning to count what’s important
Golovneva, O., Wang, T., Weston, J., and Sukhbaatar, S · 2024
Closest in time.
Rotary position embedding for vision transformer
Heo, B., Park, S., Han, D., and Yun, S · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Lie group algebra convolutional filters
Kumar, H., Parada-Mayorga, A., and Ribeiro, A · 2024
Closest in time.
Vmamba: Visual state space model
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Jiao, J., and Liu, Y · 2024
Closest in time.
Transformers can do arithmetic with the right embeddings
McLeish, S., Bansal, A., Stein, A., Jain, N., Kirchenbauer, J., Bartoldson, B. R., Kailkhura, B., Bhatele, A., Geiping, J., Schwarzschild, A., et al · 2024
Closest in time.
Causal language modeling can elicit search and reasoning capabilities on logic puzzles
Shah, K., Dikkala, N., Wang, X., and Panigrahy, R · 2024
Closest in time.
Focused transformer: Contrastive training for context scaling
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2024
Closest in time.
vision-transformers-cifar10: Training vision transformers (vit) and related models on cifar-10
Yoshioka, K · 2024
Closest in time.