Fetching the paper…
Reading the bibliography…
Self-attention, as the key block of transformers, is a powerful mechanism for extracting features from the inputs.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 1906
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Transforming auto-encoders
Hinton, G. E., Krizhevsky, A., and Wang, S. D · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et al · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Rush, A. M., Chopra, S., and Weston, J · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Yu, F. and Koltun, V · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al · 2016
Earlier work this paper cites.
Joint unsupervised learning of deep representations and image clusters
Yang, J., Parikh, D., and Batra, D · 2016
Earlier work this paper cites.
Deformable convolutional networks
Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Towards k-means-friendly spaces: Simultaneous deep learning and clustering
Yang, B., Fu, X., Sidiropoulos, N. D., and Hong, M · 2017
Cited alongside, same era.
Clustering with deep learning: Taxonomy and new methods
Aljalbout, E., Golkov, V., Siddiqui, Y., Strobel, M., and Cremers, D · 2018
Cited alongside, same era.
Learning region features for object detection
Gu, J., Hu, H., Wang, L., Wei, Y., and Dai, J · 2018
Cited alongside, same era.
Matrix capsules with em routing
Hinton, G. E., Sabour, S., and Frosst, N · 2018
Cited alongside, same era.
Learning to cluster in order to transfer across domains and tasks
Hsu, Y.-C., Lv, Z., and Kira, Z · 2018
3d point capsule networks
Zhao, Y., Birdal, T., Deng, H., and Tombari, F · 2019
Later among the works it cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination
Goyal, S., Choudhury, A. R., Raje, S., Chakaravarthy, V., Sabharwal, Y., and Verma, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient attention: Attention with linear complexities
Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H · 2018
Cited alongside, same era.
An optimization view on dynamic routing between capsules
Wang, D. and Liu, Q · 2018
Cited alongside, same era.
Improving the transformer translation model with document-level context
Zhang, J., Luan, H., Sun, M., Zhai, F., Xu, J., Zhang, M., and Liu, Y · 2018
Cited alongside, same era.
Lip: Local importance-based pooling
Gao, Z., Wang, L., and Wu, G · 2019
Cited alongside, same era.
Language modeling with deep transformers
Irie, K., Zeyer, A., Schlüter, R., and Ney, H · 2019
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2019
Cited alongside, same era.
Randla-net: Efficient semantic segmentation of large-scale point clouds
Hu, Q., Yang, B., Xie, L., Rosa, S., Guo, Y., Wang, Z., Trigoni, N., and Markham, A · 2020
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Adaptive hierarchical down-sampling for point cloud classification
Nezhadarya, E., Taghavi, E., Razani, R., Liu, B., and Luo, J · 2020
Later among the works it cites.
Hopfield networks is all you need
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Gruber, L., Holzleitner, M., Pavlović, M., Sandve, G. K., Greiff, V., et al · 2020
Later among the works it cites.
Deeper or wider networks of point clouds with self-attention?
Ran, H. and Lu, L · 2020
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Later among the works it cites.
Training data-efficient image transformers and distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.