Fetching the paper…
Reading the bibliography…
The Transformer architecture consists of self-attention and feed-forward networks (FFNs) which can be viewed as key-value memories according to previous works.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., et al · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Szlam, A., Weston, J., and Fergus, R · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, A. and Astudillo, R · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Achieving human parity on automatic chinese to english news translation
Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W., Li, M., et al · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Augmenting self-attention with persistent memory
Sukhbaatar, S., Grave, E., Lample, G., Jegou, H., and Joulin, A · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2020
Later among the works it cites.
Towards making the most of context in neural machine translation
Zheng, Z., Yue, X., Huang, S., Chen, J., and Birch, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Ethayarajh, K · 2019
Cited alongside, same era.
Non-autoregressive neural machine translation with enhanced decoder input
Guo, J., Tan, X., He, D., Qin, T., Xu, L., and Liu, T.-Y · 2019
Cited alongside, same era.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
Cited alongside, same era.
Selective attention for context-aware neural machine translation
Maruf, S., Martins, A. F., and Haffari, G · 2019
Cited alongside, same era.
Sparse sequence-to-sequence models
Peters, B., Niculae, V., and Martins, A. F · 2019
Cited alongside, same era.
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Sparse attention with linear units
Zhang, B., Titov, I., and Sennrich, R · 2021
Later among the works it cites.
A length-extrapolatable transformer
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2022
Later among the works it cites.
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Later among the works it cites.
Kformer: Knowledge injection in transformer feed-forward layers
Yao, Y., Huang, S., Dong, L., Wei, F., Chen, H., and Zhang, N · 2022
Later among the works it cites.