Fetching the paper…
Reading the bibliography…
Transformers are the mainstream of NLP applications and are becoming increasingly popular in other domains such as Computer Vision.
The acl anthology network corpus
Radev, D., Muthukrishnan, P., and Qazvinian, V · 2009
Earlier work this paper cites.
Cusparse library
Naumov, M., Chien, L., Vandermersch, P., and Kapasi, U · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2011
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2012
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Snapea: Predictive early activation for reducing computation in deep convolutional neural networks
Aklaghi, V., Yazdanbakhsh, A., Samadi, K., Esmaeilzadeh, H., and Gupta, R · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Earlier work this paper cites.
Image transformer
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D · 2018
Earlier work this paper cites.
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network
Sharma, H., Park, J., Suda, N., Lai, L., Chau, B., Chandra, V., and Esmaeilzadeh, H · 2018
Earlier work this paper cites.
Prediction based execution on deep neural networks
Song, M., Zhao, J., Hu, Y., Zhang, J., and Li, T · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers, 2019
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context, 2019
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Sparten: A sparse tensor accelerator for convolutional neural networks
Gondimalla, A., Chesnut, N., Thottethodi, M., and Vijaykumar, T · 2019
Cited alongside, same era.
Efficient training of BERT by progressively stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Cited alongside, same era.
Mnnfast: A fast and scalable system architecture for memory-augmented neural networks
Jang, H., Kim, J., Jo, J.-E., Lee, J., and Kim, J · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding, 2020
Qiu, J., Ma, H., Levy, O., tau Yih, S. W., Wang, S., and Tang, J · 2020
Later among the works it cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
Efficient tensor core-based gpu kernels for structured sparsity under reduced precision
Chen, Z., Qu, Z., Liu, L., Ding, Y., and Xie, Y · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Sparse gpu kernels for deep learning
Gale, T., Zaharia, M., Young, C., and Elsen, E · 2020
Cited alongside, same era.
A3̂: Accelerating attention mechanisms in neural networks with approximation
Ham, T. J., Jung, S. J., Kim, S., Oh, Y. H., Park, Y., Song, Y., Park, J.-H., Lee, S., Park, K., Lee, J. W., and Jeong, D.-K · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Cited alongside, same era.
Duet: Boosting deep neural network efficiency on dual-module architecture
Liu, L., Qu, Z., Deng, L., Tu, F., Li, S., Hu, X., Gu, Z., Ding, Y., and Xie, Y · 2020
Cited alongside, same era.
Rethinking attention with performers, 2021
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., and Weller, A · 2021
Closest in time.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Ham, T. J., Lee, Y., Seo, S. H., Kim, S., Choi, H., Jung, S. J., and Lee, J. W · 2021
Closest in time.
Attacc the quadratic bottleneck of attention layers
Kao, S.-C., Subramanian, S., Agrawal, G., and Krishna, T · 2021
Closest in time.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2021
Closest in time.
Sparsebert: Rethinking the importance analysis in self-attention
Shi, H., Gao, J., Ren, X., Xu, H., Liang, X., Li, Z., and Kwok, J. T · 2021
Closest in time.
Neurometer: An integrated power, area, and timing modeling framework for machine learning accelerators industry track paper
Tang, T., Li, S., Nai, L., Jouppi, N., and Xie, Y · 2021
Closest in time.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2021
Closest in time.