Fetching the paper…
Reading the bibliography…
We propose Sparse Sinkhorn Attention, a new efficient and sparse method for learning to attend.
A relationship between arbitrary positive matrices and doubly stochastic matrices
Sinkhorn, R · 1964
Earlier work this paper cites.
Ranking via sinkhorn propagation
Adams, R. P. and Zemel, R. S · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, A. and Astudillo, R · 2016
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2017
Cited alongside, same era.
Tensor2tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J · 2018
Later among the works it cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2018
Cited alongside, same era.
Learning latent permutations with gumbel-sinkhorn networks
Mena, G., Belanger, D., Linderman, S., and Snoek, J · 2018
Cited alongside, same era.
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, Ł., Shazeer, N., Ku, A., and Tran, D · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Reinforced self-attention network: a hybrid of hard and soft attention for sequence modeling
Shen, T., Zhou, T., Long, G., Jiang, J., Wang, S., and Zhang, C
Cited in the paper.
Bi-directional block self-attention for fast and memory-efficient sequence modeling
Shen, T., Zhou, T., Long, G., Jiang, J., and Zhang, C
Cited in the paper.
So, D. R., Liang, C., and Le, Q. V · 2019
Later among the works it cites.
Tay, Y., Wang, S., Tuan, L. A., Fu, J., Phan, M. C., Yuan, X., Rao, J., Hui, S. C., and Zhang, A · 2019
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Closest in time.
Blockwise self-attention for long document understanding, 2020
Qiu, J., Ma, H., Levy, O., tau Yih, S. W., Wang, S., and Tang, J · 2020
Closest in time.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Closest in time.