Fetching the paper…
Reading the bibliography…
Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences.
The design for the wall street journal-based csr corpus
Paul, D. B. and Baker, J. M · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Classes for fast maximum entropy training
Goodman, J · 2001
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, F. and Bengio, Y · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
A scalable hierarchical distributed language model
Mnih, A. and Hinton, G. E · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Adaptive sampled softmax with kernel based sampling
Blanc, G. and Rendle, S · 2017
Earlier work this paper cites.
Salimans, T., Karpathy, A., Chen, X., and Kingma, D. P · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., , and Sutskever, I · 2018
Cited alongside, same era.
Self-attentional acoustic models
Sperber, M., Niehues, J., Neubig, G., Stüker, S., and Waibel, A · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Sampled softmax with random fourier features
Rawat, A. S., Chen, J., Yu, F. X. X., Suresh, A. T., and Kumar, S · 2019
Later among the works it cites.
MASS: Masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M. A., Denoyer, L., and Jegou, H · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Yang, Z., Dai, Z., Yang, Y., Carbonell, J. G., Salakhutdinov, R., and Le, Q. V · 2019
Later among the works it cites.
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Later among the works it cites.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Closest in time.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Closest in time.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Closest in time.
RoBERTa: A robustly optimized BERT pretraining approach, 2020
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2020
Closest in time.
Efficient attention: Attention with linear complexities
Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H · 2020
Closest in time.