Fetching the paper…
Reading the bibliography…
Transformer has achieved remarkable success in language, image, and speech processing.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Pearson correlation coefficient
Benesty, J., Chen, J., Huang, Y., and Cohen, I · 2009
Earlier work this paper cites.
A new iterative method for finding approximate inverses of complex matrices
Razavi, M. K., Kerayechian, A., Gachpazan, M., and Shateyi, S · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
The lj speech dataset
Ito, K · 2017
Earlier work this paper cites.
A STRUCTURED SELF-ATTENTIVE SENTENCE EMBEDDING
Lin, Z., Feng, M., dos Santos, C. N., Yu, M., Xiang, B., Zhou, B., and Bengio, Y · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R · 2018
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., and Lehtinen, J · 2018
Earlier work this paper cites.
Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms
Shen, D., Wang, G., Wang, W., Min, M. R., Su, Q., Zhang, Y., Li, C., Henao, R., and Carin, L · 2018
Earlier work this paper cites.
Accelerating neural transformer via an average attention network
Zhang, B., Xiong, D., and Su, J · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
Fabbri, A., Li, I., She, T., Li, S., and Radev, D · 2019
Earlier work this paper cites.
A review on deep learning techniques for 3d sensed data classification
Griffiths, D. and Boehm, J · 2019
Earlier work this paper cites.
Star-transformer
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z · 2019
Earlier work this paper cites.
Low-rank and locality constrained self-attention for sequence modeling
Guo, Q., Qiu, X., Xue, X., and Zhang, Z · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Schneider, S., Baevski, A., Collobert, R., and Auli, M · 2019
Earlier work this paper cites.
What do single-view 3d reconstruction networks learn?
Tatarchenko, M., Richter, S. R., Ranftl, R., Li, Z., Koltun, V., and Brox, T · 2019
Cited alongside, same era.
ETC: Encoding long and structured inputs in transformers
Ainslie, J., Ontanon, S., Alberti, C., Cvicek, V., Fisher, Z., Pham, P., Ravula, A., Sanghai, S., Wang, Q., and Yang, L · 2020
Cited alongside, same era.
ETC: Encoding long and structured inputs in transformers
Ainslie, J., Ontanon, S., Alberti, C., Cvicek, V., Fisher, Z., Pham, P., Ravula, A., Sanghai, S., Wang, Q., and Yang, L · 2020
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Later among the works it cites.
cosformer: Rethinking softmax in attention
Qin, Z., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y · 2021
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Later among the works it cites.
Efficient attention: Attention with linear complexities
Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Compressed self-attention for deep metric learning with low-rank approximation
Chen, Z., Gong, M., Ge, L., and Du, B · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Dai, Z., Lai, G., Yang, Y., and Le, Q · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Evaluating the state-of-the-art of end-to-end natural language generation: The e2e nlg challenge
Dušek, O., Novikova, J., and Rieser, V · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Cited alongside, same era.
Pf-net: Point fractal network for 3d point cloud completion
Huang, Z., Yu, Y., Xu, J., Ni, F., and Le, X · 2020
Cited alongside, same era.
Synthesizer: Rethinking self-attention for transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C · 2021
Later among the works it cites.
Cluster-former: Clustering-based sparse transformer for question answering
Wang, S., Zhou, L., Gan, Z., Chen, Y.-C., Fang, Y., Sun, S., Cheng, Y., and Liu, J · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Later among the works it cites.
Lazyformer: Self attention with lazy update
Ying, C., Ke, G., He, D., and Liu, T.-Y · 2021
Later among the works it cites.
Pointr: Diverse point cloud completion with geometry-aware transformers
Yu, X., Rao, Y., Wang, Z., Liu, Z., Lu, J., and Zhou, J · 2021
Later among the works it cites.
You only sample (almost) once: Linear cost self-attention via bernoulli sampling
Zeng, Z., Xiong, Y., Ravi, S., Acharya, S., Fung, G. M., and Singh, V · 2021
Later among the works it cites.
Poolingformer: Long document modeling with pooling attention
Zhang, H., Gong, Y., Shen, Y., Li, W., Lv, J., Duan, N., and Chen, W · 2021
Later among the works it cites.
Long-short transformer: Efficient transformers for language and vision
Zhu, C., Ping, W., Xiao, C., Shoeybi, M., Goldstein, T., Anandkumar, A., and Catanzaro, B · 2021
Later among the works it cites.
Proteinbert: A universal deep-learning model of protein sequence and function
Brandes, N., Ofer, D., Peleg, Y., Rappoport, N., and Linial, M · 2022
Closest in time.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Closest in time.
Paragen: A parallel generation toolkit
Feng, J., Zhou, Y., Zhang, J., Qian, X., Wu, L., Zhang, Z., Liu, Y., Wang, M., Li, L., and Zhou, H · 2022
Closest in time.
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J · 2022
Closest in time.
Transformer quality in linear time
Hua, W., Dai, Z., Liu, H., and Le, Q · 2022
Closest in time.
Efficient long-text understanding with short-text models
Ivgi, M., Shaham, U., and Berant, J · 2022
Closest in time.
FNet: Mixing tokens with Fourier transforms
Lee-Thorp, J., Ainslie, J., Eckstein, I., and Ontanon, S · 2022
Closest in time.
ABC: Attention with bounded-memory control
Peng, H., Kasai, J., Pappas, N., Yogatama, D., Wu, Z., Kong, L., Schwartz, R., and Smith, N. A · 2022
Closest in time.
ABC: Attention with bounded-memory control
Peng, H., Kasai, J., Pappas, N., Yogatama, D., Wu, Z., Kong, L., Schwartz, R., and Smith, N. A · 2022
Closest in time.
cosformer: Rethinking softmax in attention
Qin, Z., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y · 2022
Closest in time.
Image super-resolution via iterative refinement
Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M · 2022
Closest in time.
Scrolls: Standardized comparison over long language sequences
Shaham, U., Segal, E., Ivgi, M., Efrat, A., Yoran, O., Haviv, A., Gupta, A., Xiong, W., Geva, M., Berant, J., et al · 2022
Closest in time.
Flowformer: Linearizing transformers with conservation flows
Wu, H., Wu, J., Xu, J., Wang, J., and Long, M · 2022
Closest in time.
Simple local attentions remain competitive for long-context tasks
Xiong, W., Oguz, B., Gupta, A., Chen, X., Liskovich, D., Levy, O., Yih, S., and Mehdad, Y · 2022
Closest in time.
Linear complexity randomized self-attention mechanism
Zheng, L., Wang, C., and Kong, L · 2022
Closest in time.