Fetching the paper…
Reading the bibliography…
Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 1904
Earlier work this paper cites.
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J · 2014
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., Fergus, R., et al · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Gruslys, A., Munos, R., Danihelka, I., Lanctot, M., and Graves, A · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y · 2016
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
Oord, A. v. d., Kalchbrenner, N., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Van Den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Efficient attention using a fixed-size memory representation
Britz, D., Guan, M. Y., and Luong, M.-T · 2017
Cited alongside, same era.
Pixelsnail: An improved autoregressive generative model
Chen, X., Mishra, N., Rohaninejad, M., and Abbeel, P · 2017
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Al-Rfou, R., Choe, D., Constant, N., Guo, M., and Jones, L · 2018
Later among the works it cites.
Transformer-xl: Language modeling with longer-term dependency
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2018
Later among the works it cites.
The challenge of realistic music generation: modelling raw audio at scale
Dieleman, S., van den Oord, A., and Simonyan, K · 2018
Later among the works it cites.
An improved relative self-attention mechanism for transformer with application to music generation
Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Hawthorne, C., Dai, A. M., Hoffman, M. D., and Eck, D · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiu, C.-C. and Raffel, C · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaev, O., Venkatesh, G., et al · 2017
Cited alongside, same era.
Parallel multiscale autoregressive density estimation
Reed, S., Oord, A. v. d., Kalchbrenner, N., Colmenarejo, S. G., Wang, Z., Belov, D., and de Freitas, N · 2017
Cited alongside, same era.
Salimans, T., Karpathy, A., Chen, X., and Kingma, D. P · 2017
Cited alongside, same era.
Generating wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N · 2018
Later among the works it cites.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
Menick, J. and Kalchbrenner, N · 2018
Later among the works it cites.
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, Ł., Shazeer, N., and Ku, A · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Later among the works it cites.